AI Engineer · New Delhi → remote

I build real-time voice AI end to end.

From VITS voice models to the full stack that serves them in production — LLM/RAG pipelines, computer vision, React · FastAPI · Postgres. Real systems, real latency, real users.

scroll

01 — About

I ship generative-AI products end-to-end — and I sweat the milliseconds that make them feel human.

I'm an AI Engineer focused on real-time, production systems: LLM/RAG pipelines, computer-vision models, and the full stack (React · FastAPI · PostgreSQL) that serves them to real people. My current focus is voice — automated interviews that listen, think, and speak back in under three seconds.

  • 0students / month, live
  • 0campuses in production
  • 0per 30-min interview
  • 0voice response latency

MIT Solve — S&P Global StepForward · 2026 global winner, 1 of 6 · $200K grant

02 — Selected work

Things I built that shipped.

  1. AI Interviewer

    NavGurukul · 2025 · flagship

    A real-time voice-interview platform in production across 9 campuses, interviewing 1000+ students a month at ~$0.04 per 30-minute session.

    • Real-time voice loop: browser STT, Piper VITS TTS, Silero VAD & barge-in — sub-3s time-to-first-audio via client-side orchestration + edge calls.
    • A JD→CEFR rubric engine that turns a job description into a skill rubric and adapts questions to probe each skill.
    • Reads candidate projects via OCR; an async scoring pipeline grades each interview into a leveled scorecard.
    • React
    • FastAPI
    • PostgreSQL
    • Cloudflare Worker
    • Piper
    • Hugging Face
    • Tesseract

    The voice engine on this very page is a slice of this system → Voice Lab

  2. Speech-to-Speech Voice Models

    NavGurukul · 2025

    Trained multi-accent voice models (Hindi, Gujarati, Hinglish) with VITS end-to-end TTS, and shipped a client-side STT pipeline at $0.12/hour.

    • Built the training-data pipeline: grapheme-to-phoneme (eSpeak) + forced alignment (MFA).
    • Trained with PyTorch Lightning on AWS SageMaker, exported to ONNX for fast, low-cost inference.
    • Client-side STT (Web Speech API) + synthesized voice out — the lineage behind this site's Voice Lab.
    • VITS
    • PyTorch Lightning
    • MFA
    • eSpeak
    • ONNX
    • SageMaker
  3. Product Lens

    Dynamic Mavens · 2024

    A vision-based product scanner that identifies products from images for a coffee / brewing-gadgets brand.

    • Automated the dataset pipeline — labeling, augmentation, training & validation on Roboflow — reaching 92% detection accuracy.
    • Image-similarity search (OpenCV + scikit-learn) matches a scanned product to the catalog.
    • Also engineered a hybrid recommender (collaborative + content-based) that lifted CTR +40%.
    • OpenCV
    • scikit-learn
    • Roboflow
    • YOLO
    • Image search

03 — Experience

  1. Jul 2025 — now

    AI / ML Engineer · NavGurukul (Remote)

    Built & scaled the AI Interviewer end-to-end. Shipped the Telangana Government's “Draw & Learn” sketch-recognition (CNNs on Quick-Draw, SageMaker), and AI learning tools used across campuses.

  2. Jan — Jun 2025

    AI / ML Intern · BlackNGreen Mobile Solutions

    Voice & chat agents for Nepal's largest telecom — multi-turn conversation, RAG context retrieval, Redis caching. Multi-agent sales workflow on Vapi. Benchmarked vector DBs & RAG evaluation (Ragas).

  3. Oct — Dec 2024

    AI / ML Intern · Dynamic Mavens

    Built the vision product scanner (Product Lens) and a hybrid recommendation model that lifted click-through rate by 40%.

  4. 2021 — 2025

    B.Tech, AI & Machine Learning · GGSIPU

    Delhi Technical Campus · 8.2 GPA.

04 — Voice Lab

Talk to my actual voice stack.

This runs the real client-side pipeline from the AI Interviewer — browser STT + Piper VITS speaking back — right here, no server, no LLM. Speak, and hear it in the voice I trained.

Best in Chrome, Edge or Safari (live transcription uses the Web Speech API). The voice model is Piper VITS running in your browser via ONNX Runtime — the same architecture behind the interviewer.

05 — Toolkit

AI / ML & GenAI

NLP · RAG · LLM fine-tuning · Prompt engineering · Transformers · STT / TTS · Model evaluation · CNN · RNN · GAN · YOLO · PyTorch · TensorFlow

Agents & frameworks

Multi-agent workflows · LangChain · LangGraph · AutoGen · Haystack · Hugging Face · Sentence Transformers

Full-stack & data

Python · FastAPI · Node.js · React · PostgreSQL · REST · WebSockets · Event-driven systems

MLOps & tooling

AWS SageMaker · MLflow · Vector databases · Redis · ETL · Jupyter · Postman · Power BI · Tableau

06 — Recognition & beyond

MIT Solve · S&P Global StepForward Challenge

2026 global winner — 1 of 6. My AI Interviewer & Speech-to-Speech work anchored NavGurukul's winning AI Learning Ecosystem, awarded a $200K grant.

  • Vice President, AWAAZ — the literature society
  • Host, GAMESCON gaming-fest
  • Top-10 finalist, National Hackathon
  • Captain, college basketball team

07 — Contact

Let's build something that talks back.

Have a role, a project, or just want to say hi? Leave a note — or .

Voice mode: say a field's text, “next field” to move on, “send it” to submit.