AI Engineer · New Delhi → remote
I build real-time voice AI end to end.
From VITS voice models to the full stack that serves them in production — LLM/RAG pipelines, computer vision, React · FastAPI · Postgres. Real systems, real latency, real users.
01 — About
I ship generative-AI products end-to-end — and I sweat the milliseconds that make them feel human.
I'm an AI Engineer focused on real-time, production systems: LLM/RAG pipelines, computer-vision models, and the full stack (React · FastAPI · PostgreSQL) that serves them to real people. My current focus is voice — automated interviews that listen, think, and speak back in under three seconds.
- 0students / month, live
- 0campuses in production
- 0per 30-min interview
- 0voice response latency
MIT Solve — S&P Global StepForward · 2026 global winner, 1 of 6 · $200K grant
02 — Selected work
Things I built that shipped.
-
AI Interviewer
NavGurukul · 2025 · flagship
A real-time voice-interview platform in production across 9 campuses, interviewing 1000+ students a month at ~$0.04 per 30-minute session.
- Real-time voice loop: browser STT, Piper VITS TTS, Silero VAD & barge-in — sub-3s time-to-first-audio via client-side orchestration + edge calls.
- A JD→CEFR rubric engine that turns a job description into a skill rubric and adapts questions to probe each skill.
- Reads candidate projects via OCR; an async scoring pipeline grades each interview into a leveled scorecard.
- React
- FastAPI
- PostgreSQL
- Cloudflare Worker
- Piper
- Hugging Face
- Tesseract
The voice engine on this very page is a slice of this system → Voice Lab
-
Speech-to-Speech Voice Models
NavGurukul · 2025
Trained multi-accent voice models (Hindi, Gujarati, Hinglish) with VITS end-to-end TTS, and shipped a client-side STT pipeline at $0.12/hour.
- Built the training-data pipeline: grapheme-to-phoneme (eSpeak) + forced alignment (MFA).
- Trained with PyTorch Lightning on AWS SageMaker, exported to ONNX for fast, low-cost inference.
- Client-side STT (Web Speech API) + synthesized voice out — the lineage behind this site's Voice Lab.
- VITS
- PyTorch Lightning
- MFA
- eSpeak
- ONNX
- SageMaker
-
Product Lens
Dynamic Mavens · 2024
A vision-based product scanner that identifies products from images for a coffee / brewing-gadgets brand.
- Automated the dataset pipeline — labeling, augmentation, training & validation on Roboflow — reaching 92% detection accuracy.
- Image-similarity search (OpenCV + scikit-learn) matches a scanned product to the catalog.
- Also engineered a hybrid recommender (collaborative + content-based) that lifted CTR +40%.
- OpenCV
- scikit-learn
- Roboflow
- YOLO
- Image search
03 — Experience
-
Jul 2025 — now
AI / ML Engineer · NavGurukul (Remote)
Built & scaled the AI Interviewer end-to-end. Shipped the Telangana Government's “Draw & Learn” sketch-recognition (CNNs on Quick-Draw, SageMaker), and AI learning tools used across campuses.
-
Jan — Jun 2025
AI / ML Intern · BlackNGreen Mobile Solutions
Voice & chat agents for Nepal's largest telecom — multi-turn conversation, RAG context retrieval, Redis caching. Multi-agent sales workflow on Vapi. Benchmarked vector DBs & RAG evaluation (Ragas).
-
Oct — Dec 2024
AI / ML Intern · Dynamic Mavens
Built the vision product scanner (Product Lens) and a hybrid recommendation model that lifted click-through rate by 40%.
-
2021 — 2025
B.Tech, AI & Machine Learning · GGSIPU
Delhi Technical Campus · 8.2 GPA.
04 — Voice Lab
Talk to my actual voice stack.
This runs the real client-side pipeline from the AI Interviewer — browser STT + Piper VITS speaking back — right here, no server, no LLM. Speak, and hear it in the voice I trained.
Best in Chrome, Edge or Safari (live transcription uses the Web Speech API). The voice model is Piper VITS running in your browser via ONNX Runtime — the same architecture behind the interviewer.
05 — Toolkit
AI / ML & GenAI
NLP · RAG · LLM fine-tuning · Prompt engineering · Transformers · STT / TTS · Model evaluation · CNN · RNN · GAN · YOLO · PyTorch · TensorFlow
Agents & frameworks
Multi-agent workflows · LangChain · LangGraph · AutoGen · Haystack · Hugging Face · Sentence Transformers
Full-stack & data
Python · FastAPI · Node.js · React · PostgreSQL · REST · WebSockets · Event-driven systems
MLOps & tooling
AWS SageMaker · MLflow · Vector databases · Redis · ETL · Jupyter · Postman · Power BI · Tableau
06 — Recognition & beyond
MIT Solve · S&P Global StepForward Challenge
2026 global winner — 1 of 6. My AI Interviewer & Speech-to-Speech work anchored NavGurukul's winning AI Learning Ecosystem, awarded a $200K grant.
- Vice President, AWAAZ — the literature society
- Host, GAMESCON gaming-fest
- Top-10 finalist, National Hackathon
- Captain, college basketball team
07 — Contact
Let's build something that talks back.
Have a role, a project, or just want to say hi? Leave a note — or .