Data Engineer · IIT Kharagpur '24

I build pipelines that move billions of records a day.

$ |

I'm Sanskar Gupta — a Data Engineer at MobileWalla building production ETL pipelines on AWS that process 10TB+ daily. I've cut cloud costs by $96K/year and taken a 42-hour Spark job down to 5 hours — because at scale, performance is a feature.

Let's connect → Download resume View projects
daas_pipeline.scala — production DAG (hover the nodes)
Kafka — real-time streaming ingestion Kafka streaming S3 — raw data lake S3 raw data lake Apache Spark on EMR — 10TB+ transformed daily, orchestrated by Airflow Spark · EMR 10TB+/day · Airflow Athena — external tables (Parquet / AVRO / CSV) Athena parquet · avro Redshift & RDS — analytics warehouse Redshift · RDS analytics Client delivery — 15+ clients onboarded Client 15+ live
0
Data processed daily
$0
Annual AWS cost savings
0
Runtime cut: 42h → 5h Spark job
0
Clients onboarded end-to-end
01 · BASICS

Engineer first, optimizer always.

I work where scale meets cost. At MobileWalla I build and run production ETL pipelines that move 10TB+ of data every day across Spark, EMR, Airflow, and Kafka — comfortable across batch processing, real-time streaming, and data governance.

My favorite kind of problem: a pipeline that's slow, expensive, or both. I've taken a 42-hour Spark job down to 5 hours by fixing a disk-spill bottleneck, designed a spot-instance strategy with 2-tier failure recovery that cut EMR costs 25%, and shipped GDPR/CCPA-compliant opt-out logic across pipelines — $96K/year saved in total.

Outside work, I write about Spark internals and ETL patterns on LinkedIn for 2,300+ followers — because explaining shuffle behavior in plain words is the best way to prove you understand it.

ScalaPythonSQLC++ Apache SparkPySparkSpark Streaming AirflowKafkaAWS EMR S3AthenaRedshift AWS GlueCloudflare R2FastAPI DjangoMySQLMongoDBTableau
NameSanskar Gupta
RoleData Engineer
CompanyMobileWalla, Kolkata
Experience2 years
EducationIIT Kharagpur '24
Open toOnsite · Hybrid · Remote (India)
Status● Open to opportunities
02 · EDUCATION

Where the fundamentals come from.

B.Tech in Electrical Engineering

Indian Institute of Technology, Kharagpur · 2020 – 2024

Strong base in data structures, algorithms, and systems thinking — the same fundamentals I now apply to query planning, shuffle optimization, and distributed system design. Competitive programmer on the side: AIR 31 in CodeKaze and a 3-star CodeChef rating.

GPA 8.32 / 10 DSA · C++ · Python CodeChef ★★★ — sanskar12k
03 · CAREER & IMPACT

What I've shipped — and what it changed.

Data Engineer

Jul 2024 — Present
MobileWalla · Kolkata — consumer-intelligence data at planetary scale
  • Built end-to-end ETL pipelines in Scala, Spark & Airflow, onboarding 15+ clients while processing 10TB+ data daily.
  • Reduced AWS EMR costs by 25% ($66K/year) with a diversified spot-instance setup, plus a 2-tier failure recovery system (step retry → on-demand cluster fallback) to protect SLAs.
  • Cut a pipeline's cost by 75% ($30K/year) — resolved a Spark disk-spill bottleneck, handled data skew, and tuned configs to take the job from 42 hours to 5.
  • Orchestrated daily & weekly Airflow DAGs computing metrics into Redshift & RDS, powering client analytics.
  • Designed a pipeline observability system tracking daily ingestion volumes with SLA-based alerting and data-quality handling.
  • Implemented GDPR/CCPA-compliant opt-out logic and IP-overcrowding rules across ETL pipelines, strengthening data governance.
ScalaSparkAirflowKafkaEMR S3AthenaRedshiftRDSGlue

Engineering Development Group (EDG) Intern

May — Jul 2023
MathWorks — makers of MATLAB & Simulink
  • Created Simulink blocks for a Pulse Oximeter and DS18B20 sensor using the IO Device Builder App, expanding MATLAB's hardware support.
  • Validated I2C sensor-creation methodology and extended it to design the L3G4200D gyroscope block; refactor cut sensor onboarding effort by 30%.
  • Discovered 5 critical bugs in the pre-release MATLAB R2023 during a company-wide bashing event.
MATLABSimulinkI2CEmbedded

Product Development Intern

May — Jul 2022
Narrato — content-creation platform
  • Revamped a collaborative creator platform into a single-user product, retaining core functionality while improving usability.
  • Built the full stack with Django, MySQL, Quill.JS & Bootstrap.
  • Devised a 3-tier plan-based access control system for dynamic, plan-aligned feature permissions.
DjangoMySQLQuill.JSBootstrap
04 · WHAT I'M LOOKING FOR

The next system I want to build.

I'm looking for teams where data is the product — and where scale, latency, and cost actually matter.

Roles

Data Engineer · Big Data Engineer · Software Engineer (Data) — with real ownership of pipelines and platforms.

Where & how

India-based — onsite, hybrid, or remote. Comfortable in fast-moving product teams and async-first setups alike.

Problems I want

Streaming & lakehouse architectures, multi-terabyte batch, cost-aware infra — ideally in adtech, fintech, or B2B SaaS.

Actively interviewing. If your team runs Spark, Kafka, or a serious AWS data stack — I'd love to talk.
05 · PROJECTS

Things I've built on my own time.

ELT · AI-powered

Insight Stream

End-to-end ELT pipeline for e-commerce market research: scrapes 500+ products & 5,000+ reviews per query via Playwright with 10 multithreaded workers (−80% collection time, with retries & rate limiting). Cloudflare R2 as the data lake, Pandas for enrichment — Bayesian-weighted product scores, brand rankings with confidence intervals, price segmentation — plus LLM-based multi-dimensional sentiment analysis on 100+ reviews per batch. FastAPI backend with MySQL and Google OAuth 2.0 delivers search-to-insight in 20 minutes.

Playwright → R2 data lake → Pandas + LLM → FastAPI / MySQL
PythonFastAPIPlaywrightPandasClaude APICloudflare R2MySQL
View on GitHub ↗
Real-time streaming

Real-Time Spam Detection Pipeline

Streaming spam detection processing 10k+ events/min with Spark Structured Streaming and Kafka. Sliding-window aggregations with watermarking flag users posting 10+ comments/min and posts with 5+ flagged users; flagged entities are routed to dedicated Kafka output topics so downstream pipelines consume pre-filtered streams.

Kafka → Spark Structured Streaming → windowed agg → output topics
ScalaSpark Structured StreamingKafkaWatermarking
View on GitHub ↗
06 · WRITING

I explain Spark so it sticks.

Technical posts on Spark internals and ETL patterns, read by 2,300+ engineers on LinkedIn.

07 · ACHIEVEMENTS

Proof of competitive edge.

1st place · Web Dev Hackathon

Spearheaded a team to win among 30+ teams at KodeInKgp.

AIR 31 · CodeKaze

National coding contest by Coding Ninjas · 3★ on CodeChef (sanskar12k).

Finalist · Dark Pattern Hackathon

Reached the finale among 875+ teams — IIT BHU × Dept. of Consumer Affairs.

08 · CONNECT

Got a data problem worth solving?

Whether it's a role, a referral, or just a conversation about Spark shuffle internals — my inbox is open.

Email copied ✓