About
Data Analyst and Data Scientist with 2+ years of experience building production ML pipelines, optimising large-scale ETL systems, and shipping analytics-driven features for live platforms. At SEBI, led the overhaul of a Hive-based ETL pipeline processing 10+ TB of daily trading data, cutting ETL runtime by ~35% and storage by ~60% through a CSV-to-Parquet migration and Spark tuning. As a freelance at CloudsAvenue, built and deployed ML models and LLM-driven recommendation features used by three live consumer platforms, and engineered the backend data architecture powering a large-scale wildlife knowledge repository. Strong foundation in Python, SQL, statistical modeling, and BI tooling (Power BI, Tableau, Excel).
Skills & Expertise (30)
Work Experience
Freelance Analytics Manager (Data & Ops)
CloudsAvenue TechSolutions LLP
Oct 2024 - Apr 2026
Developed and integrated custom ML models for wildlife sighting prediction and LLM-driven recommendation features, connecting Python-based model outputs to live, user-facing interfaces across three production platforms. Implemented end-to-end data pipelines linking ML backends to interactive frontend dashboards, ensuring reliable, low-latency model serving in production. Engineered the backend data architecture for a large-scale wildlife knowledge repository (Explore Jungles), supporting complex querying across a multi-category species and location dataset. Built and deployed three client-facing platforms (Junglore, A9i Innovations, Grand Festival Mumbai) end-to-end, owning the full stack from database architecture to production deployment using the MERN stack and Next.js. Optimised server-side API performance across all platforms, reducing inter-service latency through async processing and endpoint caching strategies.
Data Analytics Intern
Securities and Exchange Board of India (SEBI)
Dec 2022 - Dec 2023
Led the overhaul of a Hive-based ETL pipeline processing 10+ TB of daily trading data — migrated storage from CSV to Parquet (~60% storage reduction), rewrote Spark scripts, and reduced ETL runtime by ~35%, ensuring SLA-compliant data delivery. Built an NLP-based address similarity engine using Sentence Transformers (all-mpnet-base-v2) via Hugging Face, generating 768-dimensional embeddings for semantic matching across customer demographic data; deployed as a production Django web app with a real-time CSV upload interface. Implemented an AutoML pipeline using H2O.ai with automated data cleaning and hyperparameter optimisation, improving predictive model accuracy by ~15% over the previous baseline. Built a real-time stock ticker extraction pipeline (PINAKA) using OCR and YOLO object detection on Zee Business livestreams, automating extraction of stock names, symbols, and recommendations — eliminating an estimated 3+ hours of daily manual analyst effort.
Sports Analyst
Hudl India Pvt. Ltd.
Aug 2021 - Nov 2021
Used video analytics to evaluate player skill levels and identify areas for improvement, streamlining the film evaluation process and improving the efficiency and accuracy of performance assessments.
Education
M.Sc Data Science - SIES College of Arts, Science and Commerce
2021 - 2023 · Afghanistan
B.Sc Computer Science - SIES College of Arts, Science and Commerce
2018 - 2021 · Afghanistan
Certifications
Microsoft Power BI for Beginners
· 2024
Applied Data Science with Python
· 2022
Tableau Training
· 2022
Interested in this developer?
Profile Score Breakdown
Profile Overview
Skills (30)
Click a skill to find developers with the same skill