I'm Sanskar Gupta — a Data Engineer at MobileWalla building production ETL pipelines on AWS that process 10TB+ daily. I've cut cloud costs by $96K/year and taken a 42-hour Spark job down to 5 hours — because at scale, performance is a feature.
I work where scale meets cost. At MobileWalla I build and run production ETL pipelines that move 10TB+ of data every day across Spark, EMR, Airflow, and Kafka — comfortable across batch processing, real-time streaming, and data governance.
My favorite kind of problem: a pipeline that's slow, expensive, or both. I've taken a 42-hour Spark job down to 5 hours by fixing a disk-spill bottleneck, designed a spot-instance strategy with 2-tier failure recovery that cut EMR costs 25%, and shipped GDPR/CCPA-compliant opt-out logic across pipelines — $96K/year saved in total.
Outside work, I write about Spark internals and ETL patterns on LinkedIn for 2,300+ followers — because explaining shuffle behavior in plain words is the best way to prove you understand it.
Strong base in data structures, algorithms, and systems thinking — the same fundamentals I now apply to query planning, shuffle optimization, and distributed system design. Competitive programmer on the side: AIR 31 in CodeKaze and a 3-star CodeChef rating.
I'm looking for teams where data is the product — and where scale, latency, and cost actually matter.
Data Engineer · Big Data Engineer · Software Engineer (Data) — with real ownership of pipelines and platforms.
India-based — onsite, hybrid, or remote. Comfortable in fast-moving product teams and async-first setups alike.
Streaming & lakehouse architectures, multi-terabyte batch, cost-aware infra — ideally in adtech, fintech, or B2B SaaS.
End-to-end ELT pipeline for e-commerce market research: scrapes 500+ products & 5,000+ reviews per query via Playwright with 10 multithreaded workers (−80% collection time, with retries & rate limiting). Cloudflare R2 as the data lake, Pandas for enrichment — Bayesian-weighted product scores, brand rankings with confidence intervals, price segmentation — plus LLM-based multi-dimensional sentiment analysis on 100+ reviews per batch. FastAPI backend with MySQL and Google OAuth 2.0 delivers search-to-insight in 20 minutes.
Streaming spam detection processing 10k+ events/min with Spark Structured Streaming and Kafka. Sliding-window aggregations with watermarking flag users posting 10+ comments/min and posts with 5+ flagged users; flagged entities are routed to dedicated Kafka output topics so downstream pipelines consume pre-filtered streams.
Technical posts on Spark internals and ETL patterns, read by 2,300+ engineers on LinkedIn.
What actually happens in the physical plan when you "just grab a few rows" — and the benchmark that surprised people.
Read on LinkedIn →Skew, spill, and shuffle-partition tuning on a real production geofencing pipeline — with the configs that mattered.
Read on LinkedIn →How diversified instance families plus retry-and-fallback orchestration cut EMR costs 25% without breaking SLAs.
Read on LinkedIn →Spearheaded a team to win among 30+ teams at KodeInKgp.
National coding contest by Coding Ninjas · 3★ on CodeChef (sanskar12k).
Reached the finale among 875+ teams — IIT BHU × Dept. of Consumer Affairs.
Whether it's a role, a referral, or just a conversation about Spark shuffle internals — my inbox is open.