Back to Developers
Abhishek

Abhishek

Data Engineer

USA 3+ yrs exp 87 ยท Excellent

About

Data Engineer with 3+ years of experience specializing in Spark/Scala data pipeline development, ETL modernization, and legacy-to-modern architecture migrations within financial services environments. Proficient in Apache Spark, Scala, PySpark, Python, SQL, Apache Airflow, AWS (S3), and Snowflake, with strong expertise in source-to-target data reconciliation, data comparison, and validation. Proven track record of modernizing legacy ETL workflows, optimizing query performance, and ensuring high pipeline reliability for critical reporting and analytics.

Skills & Expertise (59)

Apache Spark Advanced
8.5/10
4
Years Exp
Python Advanced
8.0/10
4
Years Exp
SQL Advanced
8.0/10
4
Years Exp
Data Pipeline Development Advanced
8.0/10
4
Years Exp
ETL Advanced
8.0/10
4
Years Exp
Scala Intermediate
7.5/10
4
Years Exp
Data Warehousing Intermediate
7.5/10
4
Years Exp
AWS Intermediate
7.5/10
4
Years Exp
Docker Intermediate
7.0/10
4
Years Exp
Git Intermediate
7.0/10
4
Years Exp
Oracle Data Modeling JDBC Parquet CSV JSON REST APIs Workflow Orchestration Apache Airflow Azure Data Lake Storage Azure Data Factory Microsoft SQL Server Postgresql MySql AWS Glue Star Schema Query Optimization partitioning GitHub LINUX Agile Scrum code reviews Pandas Java Advanced SQL Bash Shell Scripting Object-Oriented Programming Spark SQL PySpark ELT Batch Processing Data Ingestion Data Transformation Data Cleaning Azure NumPy Data Validation Data Comparison Schema Validation Error Handling Retry Mechanisms Logging Monitoring Amazon S3 AWS Lambda Amazon Redshift Snowflake

Work Experience

Data Engineer

State Street

Sep 2025 - Present

Engineered Spark and Scala data-ingestion workflows for large-scale financial datasets, applying incremental loading and validation checks to process approximately 250,000 records daily while modernizing legacy pipeline architecture. Modernized and optimized Snowflake and Spark SQL transformations through query restructuring, selective column retrieval, and warehouse-aware processing, reducing representative analytical query execution time by 28% across recurring reporting workloads. Automated recurring ETL data pipeline workflows with Apache Airflow and Scala jobs, configuring task dependencies, retries, and failure notifications to reduce manual pipeline intervention from 5 to 2 hours per week. Improved data reliability for downstream financial reporting by implementing rigorous source-to-target data comparison, validation, and reconciliation checks, reducing recurring data-quality exceptions by 22% during monitored reporting cycles. Integrated AWS S3-based ingestion and Spark data-processing workflows with target data warehouse tables, coordinating across 3 source systems to standardize delivery and legacy-to-modern architecture migration requirements. Documented pipeline dependencies, data mappings, schema validation rules, and operational procedures in Git-based workflows, shortening recurring production issue triage and reconciliation debugging by 20 minutes per incident.

Data Engineer

Hexaware Technologies

Aug 2021 - Dec 2023

Developed Scala, Apache Spark, and SQL ETL pipelines to extract, transform, and load structured data from legacy relational sources into cloud data warehouses, processing 100,000 records per scheduled batch with reusable validation logic. Executed ETL modernization by implementing PySpark and Scala DataFrame transformations for large datasets, applying partition-aware processing and optimized joins to reduce representative batch execution time by 25%. Built reusable Spark and SQL transformations using CTEs, window functions, and aggregations to facilitate data comparison between legacy and modern environments, reducing manual data-preparation effort by 6 hours per reporting cycle. Strengthened data pipeline reliability through automated schema validation, null checks, duplicate detection, and source-to-target data reconciliation, reducing downstream data discrepancies by 18% during migration testing. Connected legacy databases and file-based data sources to modernized processing workflows using Spark and Snowflake Stages, coordinating across 2 delivery workstreams to resolve data mapping and architecture migration issues. Automated routine data pipeline execution and operational checks using Apache Spark tasks, Python, and Bash scripts, reducing repetitive execution steps by 30% and improving consistency during legacy-to-modern migration runs.

Education

M.S. in Data Analytics Engineering - George Mason University

- 2025 ยท Afghanistan

B.S. in Data Science & Artificial Intelligence - ICFAI Foundation for Higher Education

- 2023 ยท Afghanistan

Certifications

No certifications added yet

Interested in this developer?

Profile Score Breakdown

๐Ÿ“ท Photo 10/10
๐Ÿ“„ Resume 10/10
๐Ÿ’ผ Job Title 10/10
โœ๏ธ Bio 10/10
๐Ÿ› ๏ธ Skills 20/20
๐ŸŽ“ Education 10/10
โฑ๏ธ Experience 12/15
๐Ÿ’ฐ Rate 0/5
๐Ÿ† Certs 0/5
โœ… Verified 5/5
Total Score 87/100

Profile Overview

Member sinceSep 2026

Availability Details

Visa Status

OPT

Relocation

Open to Relocation