Soumy Pratap
Data Engineer
About
No bio added yet
Skills & Expertise (33)
Work Experience
Data Engineer
PriceEasy
Jul 2026 - Present
Engineered SQL-based ETL/ELT pipelines on Snowflake, processing 20 GB/TB of data daily across staging and production layers, supporting 12 downstream tables/dashboards through ingestion, transformation, validation, QA, and feature engineering. Redesigned and optimized 3 existing data pipelines, cutting average runtime from 5 hrs to 2 hrs through job scheduling, orchestration, and dependency management improvements. Built AWS-based data workflows (S3, EC2, Lambda, CloudWatch) running 5 scheduled jobs, reducing pipeline failure detection time from 30 mins to 10 mins via automated monitoring and alerting. Converted TFLite models to ONNX for speaker-detection inference, reducing inference time from 80 ms to 60 ms per request. Migrated a Node.js server to Python/FastAPI, re-implementing 8 REST endpoints and cutting average API response latency from 100 ms to 80 ms.
Data Scientist
Zeta
Jul 2025 - Jan 2026
Developed a prompt optimization evaluation pipeline using TextGrad with an integrated LLM-as-critique-judge loop, iteratively refining a global system prompt and boosting generative AI output quality 2x. Designed a chatbot tool-call evaluation pipeline from scratch, built on Rasa Pro's orchestration system with predefined guardrail metrics, to benchmark decision correctness and reliability across multi-turn conversations. Performed EDA on a 2M+ row banking transaction dataset, built a population-level simulation pipeline to evaluate a new credit card product over a 1-year period, using AWS S3, Athena for scalable storage and querying, and Polars for performance. Conducted analysis on 10M+ banking transactions, leveraging PyArrow and Glue tables for smoother querying across Parquet files, and FAISS for similarity-based classification to develop income validation policies, increasing coverage by 20%. Retrained and hyperparameter-tuned an XGBoost model on a revised dataset, applying SHAP and Boruta for feature selection, PCA for dimensionality reduction, and REMO-based analysis to uncover 5+ new predictive features.
Software Development Intern
Multivirt India Pvt Ltd
May 2024 - Jul 2024
Engineered VOD platform featuring asynchronous UI updates via AJAX and dynamically rendered media assets. Integrated custom HTML5 video and Web Audio API players alongside third-party RESTful APIs for real-time metadata display. Optimized core web vitals by implementing native/observer-based lazy loading and Next-Gen image compression (WebP/AVIF), cutting initial page load times and reducing bandwidth consumption. Configured Apache HTTP Server on CentOS to act as a reverse proxy for a Node.js/Express microservice, managing backend routing and securing local API endpoints. Engineered a Node.js RESTful API integrated with a SQL database to run incoming requests through a Convolutional Neural Network (CNN) model, serving structured JSON predictions with low latency. Engineered a preprocessing pipeline using OpenCV contour analysis and Canny edge detection, accelerated via CUDA GPU parallelization to boost inference throughput; increased classification Accuracy by 10%, F1-Score by 20%, and Recall by 10%. Refactored client-side codebase and optimized asset delivery to minimize runtime memory overhead reduced Largest Contentful Paint (LCP) and First Contentful Paint (FCP) Core Web Vitals, propelling search visibility to Page 1 rankings.
Software Developer Intern
Enveu
May 2024 - Jun 2024
Engineered an enterprise observability platform utilizing FastAPI microservices to ingest, correlate, and visualize real-time AWS CloudWatch telemetry and distributed API application logs. Designed a single, strongly-typed GraphQL API endpoint serving as an abstraction layer across Elasticsearch (log analytics) and Prometheus (time-series metrics), streamlining complex query pipelines across 200,000+ log entries to reduce network overhead and over-fetching. Implemented OAuth 2.0 / JWT authentication mechanisms for RBAC-secured user access, persisting user profiles, system state, and active sessions in a cloud-hosted MongoDB cluster.
ML Research Intern
Prof Sudebkumar P Pal
May 2023 - Jul 2023
Benchmarked classical time-series (ARIMA, Prophet) and deep learning (LSTM, XGBoost) models on 10+ features, optimizing hyperparameters to minimize error metrics to MAE: 0.23 and RMSE: 0.27. Developed temporal signals such as autoregressive lag terms, rolling window statistics, and expanding averages yielding a 10% improvement in model accuracy. Implemented data-cleaning pipelines for chronological alignment, NaN interpolation, and weekly resampling; performed multivariate correlation and trend analysis via Seaborn and Matplotlib.
Education
Integrated Bachelors and Masters - Indian Institute of Technology Kharagpur
- 2025 · Afghanistan
CBSE Higher Secondary - St. Anthony’s Sr Sec School Barabanki
- 2020 · Afghanistan
Certifications
No certifications added yet
Interested in this developer?
Profile Score Breakdown
Profile Overview
Skills (33)
Click a skill to find developers with the same skill