About
Data professional experienced in building and validating predictive models (XGBoost, LSTM, SARIMAX) that reached ~85% accuracy and lifted planning outcomes by 18%. Skilled in Python, SQL, and statistical modeling, with hands-on use of AI tools (LangChain, OpenAI API, OCR) to verify AI-generated outputs against source data and support quality assurance. Comfortable translating findings into dashboards and documentation, and quick to pick up unfamiliar domains, including regulatory frameworks (ESG/BRSR/GRI/IFRS). Research background includes a peer-reviewed publication and a research internship at NIT Trichy.
Skills & Expertise (43)
Work Experience
Jr. ML Engineer
TensorixAI
May 2026 - Present
Evaluated multiple computer vision approaches (YOLOv11 object detection vs. CLIP-based embedding retrieval) for industrial part identification; selected an embedding-based method after assessing dataset constraints and recorded the trade-off analysis for the team. Built an image-embedding database using OpenAI CLIP (ViT-B/32) and PyTorch, enabling cosine-similarity-based part identification with confidence scoring for human-in-the-loop verification. Structured and labeled a dataset of 38+ industrial parts, mapping real part images to engineering drawings and converting multi-page PDF drawings into ML-ready image sets. Created a visual prediction interface surfacing predicted part number, confidence score, and top matches; outlined a phased inspection workflow (capture → identification → confidence-based validation → inventory integration). Developed a real-time supplier price-intelligence web application (Python, Streamlit, Selenium, BeautifulSoup, SQLAlchemy) automating supplier discovery and quotation extraction across multiple regions. Implemented data-extraction and validation logic (Regex, duplicate detection) to parse and verify supplier pricing, ratings, and business details; normalized pricing units (₹/Kg to ₹/Tonne) for consistent comparison. Designed a supplier-comparison engine and PostgreSQL/SQLite schema surfacing average price, cheapest supplier, and estimated procurement savings to support data-driven purchasing decisions. Architected an on-premise conversational AI assistant enabling natural-language voice and text queries against a relational database, removing the need for end users to write SQL. Designed a multi-stage NLP pipeline — automatic language detection across 6 languages followed by intent classification — and implemented a natural-language-to-SQL layer using a locally hosted LLM (Qwen2.5 via Ollama) with strict read-only query validation. Integrated speech and voice components (faster-whisper, Edge-TTS, OpenVoice V2) into a responsive FastAPI + JavaScript frontend with real-time transcription, and optimized LLM inference for constrained GPU environments through per-language routing.
Software Engineer
Zentere Solutions Private Limited
Feb 2025 - Feb 2026
Architected an end-to-end demand-forecasting system integrating external APIs (Open-Meteo) with time-series models (Prophet, XGBoost, LSTM, SARIMAX), lifting planning accuracy by 18% and reducing errors by 15–20%. Engineered lag, seasonal, and external features, and validated model outputs using MAE, RMSE, MAPE, and R² to reinforce accuracy and quality control. Built an AI-assisted analytics platform (HSense Copilot) with NLP-based querying using LangChain, enabling natural-language interaction with data and human review of AI outputs. Executed intelligent query routing and OAuth2 API integrations, and formalized an ROI engine delivering 200% ROI with a 6-month payback. Verified data outputs and flagged consumption anomalies (>10% loss), strengthening operational efficiency by 15%. Assembled ML classification models (Random Forest, KNN) for NAICS market-classification prediction (+15% precision) and built a Scope 3 emissions pipeline (+22% accuracy). Refined data quality through fuzzy matching (+20%) and constructed FastAPI + SQLite pipelines, cutting latency by 30%. Delivered ESG research and analytics workflows (BRSR, GRI, IFRS) with supporting documentation and dashboards, raising KPI-tracking efficiency by 25% and planning efficiency by 20%. Devised an AI-based document-processing system using OCR (Tesseract.js) and OpenAI pipelines, boosting extraction accuracy by 20% and processing efficiency by 30%. Produced interactive dashboards (React.js, Plotly, Matplotlib) to communicate research findings clearly, cutting analysis time by 25%.
Research Intern
NIT Trichy
May 2024 - Jul 2024
Conducted independent research implementing CNN models for hyperspectral image classification (Indian Pines, Salinas, Pavia datasets), achieving 90.2% accuracy. Applied PCA and t-SNE dimensionality-reduction techniques, cutting computation cost by 20% while improving model performance.
Data Science Intern
Prodigy Infotech
Feb 2024 - Mar 2024
Performed exploratory data analysis (EDA) and preprocessing on traffic datasets, improving analysis efficiency by 30%. Trained Decision Tree models for risk prediction, improving accuracy by 12%.
Education
Master's Course Program in Big Data Analytics - CDAC, Bengaluru
2024 - 2025 · Afghanistan
MSc Data Science - Christ University, Bengaluru
2023 - 2025 · Afghanistan
B.Sc. Statistics - St. Xavier's College, Ahmedabad
2020 - 2023 · Afghanistan
Certifications
No certifications added yet
Interested in this developer?
Profile Score Breakdown
Profile Overview
Skills (43)
Click a skill to find developers with the same skill