The Brief

a quick scan for recruiters & humans.

The Evidence

a broadcast tour of how i build.

AAWHAN VYAS

ai/ml engineer · evidence over vibes

ai · mlops · grounded rag · llm evals

AAWHAN VYAS.

ai/ml engineer   bhopal, india

AI/ML engineer who ships with receipts. Model registries, grounded RAG, and the evals that keep LLM systems honest — finished to the point where they're boring to operate. I measure everything I ship.

platforms

ML infrastructure that runs itself

ModelDock — registry, gated promotion, serving, drift watch.

grounded rag

answers that cite their source

Dasaiko — hybrid retrieval, page-level citations, live in prod.

evals

measuring what actually matters

TraceBack — deterministic ground truth vs confident nonsense.

about

I like the unglamorous parts of AI — and I like them proven.

I'm a 4th-year CS undergrad at VIT Bhopal. Most of what I build starts the same way: someone says an AI system works, and I ask against what ground truth? That question turned into TraceBack. It turned Dasaiko from a chatbot into a retrieval system with a benchmark suite. It made ModelDock refuse to promote a model that didn't earn it.

I read the papers, then write them up on LinkedIn so they stick.

Off the clock it's two disciplines that keep each other honest. Boxing — the sweet science, where you learn in real time what happens when you overcommit on vibes. And music — I produce and release as aawhan, lo-fi confessionals for the late hours; the Joji-to-Yung-Lean school of feelings. Dilruba, ionknow//idk, Aloof — out on Spotify and everywhere else.

experience

Web Developer · GenWe Films

mar 2026 — may 2026

The company's first web presence, taken from zero to signed-off production.

  • Designed and shipped genwefilms.com — React, TypeScript, Supabase, EmailJS, Vercel.
  • Owned development, deployment, stakeholder communication and final handover.
  • Incorporated 10+ rounds of client feedback into the signed-off release.
reacttypescriptsupabasevercelemailjs

ML & AI Training · Beeskilled

sep 2026 — oct 2026

Structured ML/AI training track — offer letter on file.

selected work // click a row for receipts

  1. ModelDock self-hostable ML model registry & serving platform ★9 ↗

    register → version → gate → serve: promotion is earned at the eval-threshold gate, not assumed. 72-test suite in CI/CD, authenticated inference, rate limiting, drift monitoring. MIT.

    fastapi · postgresql · redis · next.js · docker

  2. Dasaiko grounded PDF question-answering workspace — live at dasaiko.dev ★5 ↗

    Hybrid retrieval (vector + BM25 + RRF + BGE reranking): +86% recall@5, +75% MRR, +84% nDCG@10 over the RRF baseline. Answers cite the exact page of the source PDF.

    react · fastapi · postgresql/pgvector · groq

  3. TraceBack LLM incident diagnosis, evaluated against ground truth ↗

    Grounded baseline: 100% pass rate and root-cause accuracy. Qwen 2.5 3B: 10% on both — at 90% average confidence. 30 runs × 3 scenarios. The gap between fluent and correct is the whole point.

    python · fastapi · mcp · ollama · next.js

  4. Faultline MCP-native incident response: investigate → root-cause → verify ★1 ↗

    Investigates production incidents from operational tools, recommends remediation, then verifies recovery before calling it fixed.

    python · mcp

  5. ClearLabel-AI auto-annotation for industrial-safety datasets ★1 ↗

    YOLOv8 + OpenCV with a dual-layer quality audit — blur detection and confidence filtering — for human-in-the-loop labeling.

    python · yolov8 · opencv

  6. PuppetGPT RAG chatbot that answers only from your PDF ★1 ↗

    LLaMA 3 on Groq, LangChain, Streamlit. No internet search — if it's not in the document, it says so.

    python · langchain · groq · streamlit

  7. gLyric chrome extension: now-playing → genius lyrics ★1 ↗

    Detects the track playing on YouTube and opens its Genius page. Small, useful, shipped.

    javascript · chrome-extension

  8. Provena open source — PR #153: policy evaluation error logging ↗

    Fixed policy-evaluation error logging and strengthened regression coverage. ✓ 498 passed · 33 skipped — verified before merge.

    python · context-governance for agentic ai

  9. HelloblueGK open source — PR #170: XML response documentation ↗

    XML response docs for HealthController endpoints in a .NET 9 aerospace engine simulation platform. ✓ .NET build + Swagger/OpenAPI verified.

    c# · .net 9 · swagger

~50 more repos ↗

stack

languages

pythonsqljavascripttypescript

backend & apis

fastapiflaskrest apispydanticnext.jsmcp

data & ai

ragpgvectorlangchainpandasnumpyyolov8ollamagroq

databases

postgresqlredissqlite

cloud & infra

dockergithub actionsci/cdvercel

on linkedin // things i'm hyped about

follow the signal ↗

Let's build something with receipts.

AI/ML engineer focused on grounded systems and rigorous evaluation. Always open to ambitious problems, a good technical conversation, or a beat to collab on.

based inbhopal, india · IST (utc+5:30)