About
AI/ML Researcher and Developer with hands-on experience in Large Language Models (LLMs), Generative AI, NLP, and Speech. Delivered impactful solutions at IIT-Delhi, IIT-BHU, and Meril R&D, with expertise in Python, Deep Learning, and GenAI ecosystems. Skilled in NLP, vector databases, LLM fine-tuning (SFT, DPO), and building scalable AI pipelines using FastAPI and Flask. Passionate about developing real-world AI applications, including medical LLMs, multilingual speech systems, and drug recommendation systems, with a focus on research-driven innovation.
Skills & Expertise (34)
Work Experience
Research Associate – AI/ML
Adivaani, MISN Lab, IIT Delhi
Aug 2026 - Sep 2026
Developed an end-to-end multilingual speech-to-speech translation pipeline for Indian and tribal languages using Whisper, IndicConformer, SraVaani, Bhili ASR, IndicTrans2, Aadivani translation APIs, and TTS systems. Engineered a Bhili forced alignment pipeline using Montreal Forced Aligner (MFA), Epitran G2P, and Kaldi-based acoustic modeling, achieving 0.0% validation WER/LER across 3,300+ unique lexical items with phone-level TextGrid alignments. Architected a high-throughput ASR dataset generation pipeline using Qwen2.5 and vLLM, generating 11,000+ domain-specific questions across 45+ cultural domains to support low-resource Bhili and Gondi speech-data collection. Deployed research prototypes as scalable FastAPI services with both Just-based local workflows and Dockerized deployment, integrating modular ASR, translation, and TTS components for real-world multilingual speech applications.
AI Engineer
Meril (Nuvo AI), Bengaluru, India
Jun 2025 - Jul 2026
Fine-tuned medical Large Language Models (LLMs) using Supervised Fine-Tuning (SFT) and Direct Preference Optimization (DPO), owning the complete pipeline including data preprocessing, dataset curation, and instruction-format design for healthcare-focused applications. Built a multilingual, real-time speech recognition and translation pipeline by fine-tuning Whisper on Gujarati and Hindi datasets, optimized for low-latency inference using Whisper Live, C2Translate, and TensorRT. Built a real-time full-duplex voice AI pipeline using Pipecat, WebRTC, Whisper, Kokoro TTS, and LLM tool calling to enable low-latency conversational experiences. Designed and integrated a speech-to-text → LLM → text-to-speech agent architecture using RabbitMQ, Kafka, Prometheus, Grafana, and OpenTelemetry for eventing, background processing, and observability. Working on AI-driven drug recommendation systems, integrating deep learning models with knowledge graphs to deliver personalized and clinically relevant treatment insights. Architected end-to-end AI pipelines for data extraction from long medical documents (200+ pages), summarization in specified formats, question answering, two-document comparison using Retrieval-Augmented Generation (RAG), data scraping, and generation of medical documents using real-time scraped data from medical websites, deploying scalable REST APIs and integrating RabbitMQ with a publisher-subscriber model for performance optimization.
AI Research Intern
Meril, Bengaluru, India
Jan 2025 - May 2025
Implemented Indian Sign Language (ISL) recognition using a Transformer-based architecture with BERT layers, leveraging MediaPipe facial and keypoint features for gesture and expression understanding. Developed a facial recognition system using embedding-based similarity matching, enabling robust multi-angle identity recognition. Led implementation of LLM guardrails, watermarking, and quantization POCs, benchmarking multiple frameworks to improve model safety, inference speed, and deployment efficiency. Designed and deployed FastAPI-based services for speech-to-text and text-to-speech applications, supporting real-time inference workflows.
Research Intern
IIT-BHU, Varanasi, India
Apr 2024 - Aug 2024
Contributed to the Metaverse Gaze Intelligence project, developing XR-accelerated AI training for immersive learning environments. Built multilingual speaker diarization systems using Whisper, Bark, and Coqui-TTS for accurate voice segmentation and identity tagging. Fine-tuned text-to-speech and speech-to-speech models for real-time translation with voice preservation across multiple Indian languages. Deployed deep learning pipelines for audio preprocessing, speaker embedding, and alignment to enhance accuracy and inference speed.
Education
Bachelor of Technology – Artificial Intelligence and Data Science - Kakinada Institute of Engineering and Technology
2021 - 2025 · Afghanistan
Certifications
No certifications added yet
Interested in this developer?
Profile Score Breakdown
Profile Overview
Skills (34)
Click a skill to find developers with the same skill