About
Data Engineer with 3+ years of experience specializing in Spark/Scala data pipeline development, ETL modernization, and legacy-to-modern architecture migrations within financial services environments. Proficient in Apache Spark, Scala, PySpark, Python, SQL, Apache Airflow, AWS (S3), and Snowflake, with strong expertise in source-to-target data reconciliation, data comparison, and validation. Proven track record of modernizing legacy ETL workflows, optimizing query performance, and ensuring high pipeline reliability for critical reporting and analytics.
Skills & Expertise (59)
Work Experience
Data Engineer
State Street
Sep 2025 - Present
Engineered Spark and Scala data-ingestion workflows for large-scale financial datasets, applying incremental loading and validation checks to process approximately 250,000 records daily while modernizing legacy pipeline architecture. Modernized and optimized Snowflake and Spark SQL transformations through query restructuring, selective column retrieval, and warehouse-aware processing, reducing representative analytical query execution time by 28% across recurring reporting workloads. Automated recurring ETL data pipeline workflows with Apache Airflow and Scala jobs, configuring task dependencies, retries, and failure notifications to reduce manual pipeline intervention from 5 to 2 hours per week. Improved data reliability for downstream financial reporting by implementing rigorous source-to-target data comparison, validation, and reconciliation checks, reducing recurring data-quality exceptions by 22% during monitored reporting cycles. Integrated AWS S3-based ingestion and Spark data-processing workflows with target data warehouse tables, coordinating across 3 source systems to standardize delivery and legacy-to-modern architecture migration requirements. Documented pipeline dependencies, data mappings, schema validation rules, and operational procedures in Git-based workflows, shortening recurring production issue triage and reconciliation debugging by 20 minutes per incident.
Data Engineer
Hexaware Technologies
Aug 2021 - Dec 2023
Developed Scala, Apache Spark, and SQL ETL pipelines to extract, transform, and load structured data from legacy relational sources into cloud data warehouses, processing 100,000 records per scheduled batch with reusable validation logic. Executed ETL modernization by implementing PySpark and Scala DataFrame transformations for large datasets, applying partition-aware processing and optimized joins to reduce representative batch execution time by 25%. Built reusable Spark and SQL transformations using CTEs, window functions, and aggregations to facilitate data comparison between legacy and modern environments, reducing manual data-preparation effort by 6 hours per reporting cycle. Strengthened data pipeline reliability through automated schema validation, null checks, duplicate detection, and source-to-target data reconciliation, reducing downstream data discrepancies by 18% during migration testing. Connected legacy databases and file-based data sources to modernized processing workflows using Spark and Snowflake Stages, coordinating across 2 delivery workstreams to resolve data mapping and architecture migration issues. Automated routine data pipeline execution and operational checks using Apache Spark tasks, Python, and Bash scripts, reducing repetitive execution steps by 30% and improving consistency during legacy-to-modern migration runs.
Education
M.S. in Data Analytics Engineering - George Mason University
- 2025 ยท Afghanistan
B.S. in Data Science & Artificial Intelligence - ICFAI Foundation for Higher Education
- 2023 ยท Afghanistan
Certifications
No certifications added yet
Interested in this developer?
Profile Score Breakdown
Profile Overview
Availability Details
Visa Status
OPT
Relocation
Open to Relocation
Skills (59)
Click a skill to find developers with the same skill