I’m an AI Engineer who builds production-grade generative AI and retrieval systems — not prototypes with no impact.
I specialize in:
- Architecting and operating production RAG pipelines for real-time semantic search
- Designing LLM routing and grounding systems with reliability and observability
- Building scalable backend APIs for AI workflows using FastAPI and async patterns
- Evaluating LLMs for cost, latency, and consistency
🔭 Current Focus: Production RAG, LLM routing & query classification, LLM evaluation metrics, multi-agent orchestration
🚀 Target Roles: AI Engineer · Generative AI Engineer · LLM / RAG Systems Engineer · Applied Machine Learning Engineer
Python · FastAPI · Async APIs · REST · Server-Sent Events (SSE) · Docker · SQL · PostgreSQL
Generative AI · Retrieval-Augmented Generation (RAG) · Semantic Search · Large Language Models (LLMs) · LLM Evaluation & Trustworthiness Scoring LangChain · LangGraph · MCP · Prompt Engineering · Embeddings & Vector Search
ETL Pipelines · ChromaDB · FAISS · Supabase · AWS (EC2, S3) · CI/CD
Streamlit · Dashboards · SHAP Interpretability
Production-grade retrieval system combining hybrid sparse + dense search across text and images, with grounded citations and real-time streaming APIs.
- Hybrid retrieval: BM25 + vector embeddings for multi-modal semantic relevance
- Grounded citation outputs for traceable answers
- FastAPI based streaming interface using SSE
- LLM benchmarking across cost, latency, and accuracy
- Async workers, caching, and scalable backend design
What I Learned
- Hybrid retrieval architecture
- Multi-modal semantic search
- Building real-time streaming APIs
- Structuring benchmarking pipelines for LLM evaluation
Framework to assess LLM trustworthiness and behavioral consistency without ground-truth labels.
- Reference-free scoring for LLM outputs
- Behavioral consistency and multi-choice evaluation logic
- Modular Python evaluation pipelines
- Designed for reliability analysis in production-like setups
What I Learned
- LLM evaluation workflows
- Reference-free reliability metrics
- Extensible Python pipelines
- Challenges in model trustworthiness
LangGraph-based CLI client coordinating multiple MCP servers for action routing and multi-tool workflows.
- Tool orchestration (notes, weather, web search, automation)
- Session state management and contextual routing
- Extensible plugin-style MCP server integration
What I Learned
- Multi-agent orchestration patterns
- Protocol integration with MCP
- Session tracking and routing logic
- Structuring agent toolchains
Unsupervised patient stratification with HDBSCAN and SHAP-powered interpretability dashboards.
- 130k+ patient record clustering
- SHAP interpretation and interactive exploration
- End-to-end pipeline from data to UI
A multimodal system detecting beehive health via vision and audio, published at IEEE CSCITA 2023.
- Integrated YOLO vision (96.2% precision) and audio classification (99.8%)
- Real-world deployment mindset
- Hackathon winning submission
- 🥇 1st Place — Voxel51 Visual AI Hackathon | GitHub
- 📜 AWS Certified Cloud Practitioner | Credly Badge
- 🔗 LinkedIn: https://linkedin.com/in/deepmehta27
- 🖥️ GitHub: https://github.com/deepmehta27
