Back to Developers
Shamraj

Shamraj

Data Engineer

Chennai, Tamil Nadu, India 80 ยท Excellent

About

Data Engineer who has spent the past year building and running production ETL/ELT pipelines on AWS. Built a medallion-architecture pipeline (raw, stage, analytics) with AWS Glue and PySpark, and separately stood up an Iceberg-based ingestion setup that moves data from Airbyte into S3 and on to ClickHouse, with dbt models scheduled through GitHub Actions on EC2. Comfortable owning the full lifecycle - ingestion, transformation, orchestration, and reporting in Amazon QuickSight and Power BI - with a solid SQL and Python foundation and a habit of catching data quality issues early.

Skills & Expertise (20)

AWS Glue Advanced
8.5/10
1
Years Exp
PySpark Advanced
8.5/10
1
Years Exp
Amazon S3 Advanced
8.0/10
1
Years Exp
SQL Advanced
8.0/10
1
Years Exp
Python Advanced
8.0/10
1
Years Exp
Git Advanced
8.0/10
1
Years Exp
DBT Advanced
8.0/10
1
Years Exp
Amazon Athena Advanced
8.0/10
1
Years Exp
ELT Intermediate
7.5/10
1
Years Exp
ETL Intermediate
7.5/10
1
Years Exp
Amazon Ec2 Intermediate
7.5/10
1
Years Exp
Big Data Intermediate
7.0/10
1
Years Exp
Apache Spark Intermediate
7.0/10
1
Years Exp
Data Warehouse Intermediate
7.0/10
1
Years Exp
Data Quality Validation Intermediate
7.0/10
1
Years Exp
GitHub Actions Intermediate
7.0/10
1
Years Exp
Amazon QuickSight Intermediate
7.0/10
1
Years Exp
Databricks Intermediate
6.5/10
1
Years Exp
Power BI Intermediate
6.5/10
1
Years Exp
Excel Intermediate
6.0/10
1
Years Exp

Work Experience

Data Engineer

AVASOFT

Present - Present

Re-architected our source tables into a medallion pipeline - raw, stage, analytics - using AWS Glue and PySpark to progressively clean and transform data as it landed in S3. Wrote the PySpark jobs in Glue that clean, standardize, and enrich data as it moves between zones. Built a QuickSight dashboard on top of Athena so the business team could pull reporting directly from the analytics layer. Built out 320+ dbt models across 15 pipelines, now running at 450+ client accounts and roughly 22TB of data, ingesting into S3 as Iceberg tables via Airbyte and transforming into ClickHouse. Cut processing overhead noticeably by moving tables from truncate-and-load to incremental load. Fixed recurring out-of-memory (OOM) failures by tuning ClickHouse configuration and rewriting query patterns for larger data volumes. Tracked down a schema-evolution bug where ClickHouse's external tables on S3 broke on source schema changes, and fixed it by reading data directly from S3 instead of through ClickHouse. Set up the GitHub Actions workflow that runs dbt models on a cron schedule from an EC2 instance.

Data Analytics Intern

ReTech

Present - Present

Analyzed service desk data and built Tableau dashboards to track response and resolution times, which helped the team spot recurring issues and make more informed operational decisions.

Education

M.Sc. Applied Data Science - SRM Institute of Science and Technology, Kattankulathur

2023 - 2025 ยท Afghanistan

B.Sc. Computer Science - Thiruthangal Nadar College, Chennai

2019 - 2022 ยท Afghanistan

Certifications

No certifications added yet

Interested in this developer?

Profile Score Breakdown

๐Ÿ“ท Photo 10/10
๐Ÿ“„ Resume 10/10
๐Ÿ’ผ Job Title 10/10
โœ๏ธ Bio 10/10
๐Ÿ› ๏ธ Skills 20/20
๐ŸŽ“ Education 10/10
โฑ๏ธ Experience 5/15
๐Ÿ’ฐ Rate 0/5
๐Ÿ† Certs 0/5
โœ… Verified 5/5
Total Score 80/100

Profile Overview

Member sinceSep 2026

Availability Details

Visa Status

Citizen

Relocation

Open to Relocation

Skills (20)

Click a skill to find developers with the same skill