About
Data Engineer with 1+ year of experience designing and developing scalable ETL/ELT pipelines, data ingestion workflows, data lake solutions, and analytics-ready datasets using Python, SQL, PySpark, Apache Spark, Databricks, AWS S3, and EMR. Hands-on with batch processing, Delta Lake, data quality, data modeling, data warehousing, Spark performance tuning, incremental processing, workflow orchestration, and cloud data platforms. Experienced in building reliable pipelines for structured and semi-structured data and delivering reusable datasets for analytics and reporting.
Skills & Expertise (34)
Work Experience
Data Engineer
2GBR Software Private Limited
May 2025 - Present
Built and maintained 10+ ETL/ELT data pipelines using PySpark, Apache Spark, Databricks, and AWS S3 for structured and semi-structured data, improving pipeline throughput by 30%. Optimized Spark and Spark SQL workloads through partitioning, broadcast joins, predicate pushdown, caching, and query tuning, reducing average job runtime by 25% from approximately 90 minutes to 65 minutes. Designed AWS S3-based data lake workflows processing 30โ50 GB daily and delivered cleansed, analytics-ready datasets for reporting and downstream data consumption. Implemented schema validation, deduplication, transformation rules, and automated data-quality checks across more than 5 million records daily, maintaining 99%+ data accuracy. Automated pipeline scheduling, dependency management, retries, monitoring, and failure handling using Apache Airflow; developed SQL/Spark SQL transformations and worked with AWS EMR, Delta Lake, Hive, and HDFS for distributed data processing.
Education
B.Tech โ Artificial Intelligence and Data Science - S.B. Jain Institute of Technology, Management & Research
2021 - 2025 ยท Afghanistan
Certifications
No certifications added yet
Interested in this developer?
Profile Score Breakdown
Profile Overview
Skills (34)
Click a skill to find developers with the same skill