Shivani Resu
Data Engineer
About
Aspiring Data Engineer with strong hands-on knowledge of Azure Data Engineering and Databricks Lakehouse technologies. Hands-on experience working with Apache Spark (PySpark, Spark SQL) for batch data processing and transformations. Practical exposure to Azure Databricks, including Spark cluster provisioning, notebook development, and job scheduling. Experienced in building ETL pipelines using Azure Data Factory (ADF v2) with full and incremental load strategies. Hands-on implementation of Delta Lake concepts including merge operations, schema enforcement, and time travel. Implemented Slowly Changing Dimensions (SCD Type 1 and Type 2) using PySpark and Delta Lake merge functionality. Developed data validation, audit logging, and exception handling frameworks in Databricks notebooks. Strong understanding of Spark Architecture including Driver, Executors, Stages, Tasks, and Spark UI. Experience processing structured and semi-structured data formats such as CSV, Parquet, and JSON. Familiar with Azure Data Lake Storage Gen2, Blob Storage, Azure SQL, Azure Synapse, and JDBC connectivity. Implemented dynamic and parameterized notebooks using widgets and dbutils commands. Knowledge of scheduling and orchestrating data pipelines using ADF triggers and Databricks Jobs. Hands-on exposure to Azure Key Vault integration for secure credential management. Good understanding of cloud-based data architecture and Lakehouse design principles.
Skills & Expertise (20)
Work Experience
Data Engineer
Academic / Training
Present - Present
Provisioned and configured Spark clusters in Azure Databricks. Created PySpark notebooks with dynamic parameters using Databricks widgets. Read and wrote data between ADLS Gen2 and Databricks using mount points. Implemented data transformations using joins, filters, aggregations, and window functions. Developed Delta Lake tables and implemented SCD Type 1 and Type 2 using merge functionality. Implemented audit logging framework to capture record counts and missing records. Added exception handling and logging in all notebooks. Developed reusable/common notebooks and invoked them using dbutils.notebook.run. Performed data cleansing including trimming columns, duplicate checks, and key validations. Used Spark SQL and magic commands for testing and validation. Monitored and troubleshot Spark jobs using Spark UI.
Data Engineer
Academic / Training
Present - Present
Created Linked Services, Datasets, and Pipelines in Azure Data Factory v2. Implemented full load and incremental load pipelines from on-premise sources to ADLS Gen2. Developed dynamic pipelines using configuration tables with Lookup and ForEach activities. Scheduled pipelines using time-based and event-based triggers. Integrated ADF pipelines with Azure Databricks notebooks for transformations. Implemented audit logging using stored procedures in ADF pipelines. Installed and configured Self-Hosted Integration Runtime for on-premise connectivity. Implemented alert notifications using Azure Logic Apps. Created master pipelines for orchestration and dependency management.
Education
Bachelor of Technology (B.Tech) - Jawaharlal Nehru Technological University (JNTU)
- ยท Afghanistan
Certifications
No certifications added yet
Interested in this developer?
Profile Score Breakdown
Profile Overview
Skills (20)
Click a skill to find developers with the same skill