Close to 8 years building AI and platform systems that businesses run on — currently delivering production LLM, RAG and agent systems for enterprise clients including Boeing, T-Mobile and BCG. Open to relocation (UK, Europe, UAE).
Each is a small, complete, tested system — clone any of them and it runs offline in one command, no API keys, CI green.
AI / LLM Serving
- llm-gateway — one API over many LLM providers: cheapest-first routing, caching, failover with circuit breaking, cost ledger
- llm-batch-sim — discrete-event simulation quantifying why continuous batching beats naive serving (6x throughput, 9x lower p95 on the demo workload)
AI Quality & Machine Learning
- rag-eval-gate — golden-set RAG evaluation as a deterministic CI gate; a change that degrades answers fails the build and names the case
- ml-train-gate — reproducible training (logistic regression from first principles), metric-regression gates on AUC/recall, generated model cards
MLOps
- drift-watch — input drift detection as a pipeline gate: PSI with baseline-quantile bins + KS distance, unseen categories treated as maximum signal
DevOps / IaC
- tf-plan-guard — policy gate over
terraform plan: blocks destroys and replaces of protected resources before the apply
SRE / Observability
- slo-burn — error-budget arithmetic + generated multiwindow burn-rate Prometheus alerts (Google SRE Workbook policy, auto-scaled to your window)
- gpu-cost-exporter — NVIDIA DCGM metrics re-exported as money: burn, waste (the idle share, in dollars), cost per 1k inferences
Each README has a "Design decisions worth arguing with" section — the trade-offs I'd defend in a review, written down.
AI / LLM — Python, FastAPI, LangChain, LangGraph, RAG, prompt engineering, golden-set evaluation; Amazon Bedrock, Anthropic Claude, OpenAI; vLLM-style serving concerns: continuous batching, KV-cache, quantization
MLOps & Data — MLflow, Databricks, Airflow, PySpark, Kafka, dbt, BigQuery
Platform — Kubernetes (AKS/EKS/GKE), Docker, Helm, Terraform, Ansible, GitOps (ArgoCD/Flux), Azure DevOps, GitHub Actions
Observability & Reliability — Prometheus, Grafana, OpenTelemetry, SLOs, on-call, incident response; GDPR / ISO 27001 audit experience
📫 khetpalharsh@gmail.com · Open to AI Engineer / MLOps / Platform roles with visa sponsorship (UK · Europe · UAE)

