About
Data analyst with a strong statistical and programming foundation (R, Python, SQL) built through a rigorous biotechnology research background. Experienced in end-to-end analytics — from database design and advanced SQL to predictive modeling and BI dashboards — across public health, finance, and product domains. Delivered a first-author research paper and an oral conference presentation; comfortable translating messy, large-scale data into decisions.
Skills & Expertise (28)
Work Experience
Project
Credit Risk Analytics Pipeline
Jul 2026 - Jul 2026
Designed and built a normalized PostgreSQL database (4 tables, enforced foreign keys) from 443K+ raw Lending Club loan records, reducing 151 raw columns to a clean, production-ready schema through systematic data profiling and cleaning. Wrote advanced SQL analyses using CTEs, window functions (NTILE, LAG, rolling averages via ROWS BETWEEN), and a reusable stored function to segment loan risk, track quarterly cohort default rates, and generate borrower risk profiles on demand. Diagnosed a query performance bottleneck using EXPLAIN ANALYZE and cut execution time by 47% (273ms → 145ms) through targeted indexing, documented as a before/after case study. Built a 3-page Power BI dashboard connected live via DirectQuery to the database, with custom DAX measures delivering portfolio-level and risk-segmented insights; validated that Lending Club's A–G risk grades correctly predict real-world default rates (6.9% → 50.8%, monotonic across grades).
Project
Waze User Churn Prediction
Jun 2026 - Jun 2026
Built and evaluated multiple classification models (Logistic Regression, Random Forest, XGBoost) on 14,999 user records to predict app churn, engineering behavioural features from raw usage data. Achieved 18.1% recall on the churn class through iterative feature engineering and hyperparameter tuning; identified key churn drivers via feature importance analysis.
Dissertation Research Intern
ICMR-National Institute of Nutrition
Jan 2026 - May 2026
Analyzed a multi-decade national survey dataset (five survey rounds spanning 1992–2021, 700+ districts) in R, applying multilevel mixed-effects regression and cohort-based statistical modeling to identify long-term population trends. Applied spatial statistical methods (Getis-Ord Gi*) using R (spdep, sf) to detect and map statistically significant regional clusters, resolving complex multi-source geographic data-matching issues. Co-authored a research manuscript (currently under peer review) and delivered an oral presentation at the 7th Asian Population Association Conference (APAC 2026), Hanoi, Vietnam.
Project
Stroke Prediction Model
Oct 2025 - Oct 2025
Built and validated a Random Forest classifier on 5,110 patient records to predict stroke risk, achieving 81.6% sensitivity and 0.89 ROC-AUC on held-out test data. Addressed severe class imbalance (4.8% positive cases) using ROSE oversampling; identified age, average glucose level, and hypertension as the top predictive risk factors via feature importance analysis.
Education
B.Tech in Biotechnology - Dr. D. Y. Patil Biotechnology and Bioinformatics Institute
2022 - 2026 · Afghanistan
Certifications
No certifications added yet
Interested in this developer?
Profile Score Breakdown
Profile Overview
Availability Details
Visa Status
Citizen
Relocation
Depends on Offer
Skills (28)
Click a skill to find developers with the same skill