AVAILABLE FOR DATA ENGINEERING ROLES

Hi, I'm Sagar.
I build data pipelines that don't break.

Senior Data Engineer with 4+ years shipping production-grade lakehouses on Azure Databricks — from greenfield Medallion architectures ingesting Kafka CDC, to migrating 600-job BFSI estates at 10× the speed.

4+
Years · Production
526M+
Records Processed
600+
ETL Jobs Migrated
10×
Query Performance
Impact

Numbers that moved the needle.

Production-grade results across BFSI, industrial manufacturing, and energy infrastructure.

Query performance gain migrating 600+ Pentaho jobs to PySpark
0M+
Sales records re-architected from Oracle PL/SQL to Databricks
0%
Drop in data quality incidents via Bronze→Silver contracts
0 hrs
Engineering hours saved via REST API notebook automation
0%
Compute cost reduction via PySpark refactor and native UDFs
0%
Faster release cycle via first version-controlled CI/CD
0 weeks → 0 days
New-source onboarding time, parameterized PySpark framework
0+
Subsidiary business units unified on one lakehouse
Stack

The toolbox.

Built around Databricks + Spark + Delta, with deep Azure and a passing knowledge of the boring-but-essential bits.

DE Data Engineering

  • PySpark
  • Spark SQL
  • Delta Lake
  • DLT / Lakeflow
  • Spark Declarative Pipelines
  • Liquid Clustering
  • Structured Streaming
  • Kafka CDC
  • ETL / ELT
  • Data Quality
  • Source→Target Recon

DB Databricks Platform

  • Unity Catalog
  • Workflows / Jobs
  • Asset Bundles (DAB)
  • Autoloader
  • Databricks REST API
  • Lakeflow Connect
  • Notebook CI/CD
  • Job orchestration

AZ Cloud & Infra

  • Microsoft Azure
  • Azure Databricks
  • Azure Data Factory
  • ADLS Gen2
  • Azure Key Vault
  • AWS S3
  • ABFS + SAS

LG Languages

  • Python
  • SQL
  • PySpark
  • PL/SQL
  • Spark SQL
  • Bash / Shell

DS Data Sources

  • SAP HANA
  • Oracle
  • MySQL
  • Kafka
  • ADLS Gen2
  • S3
  • REST APIs

CI CI/CD & DevOps

  • Azure DevOps
  • Git
  • ARM Templates
  • DAB
  • Multi-env promotion
  • VS Code
How I build

One lakehouse, three layers, zero hand-offs.

The Medallion architecture — Bronze holds raw, Silver holds truth, Gold serves the business. Parameterized ingestion, idempotent transforms, CDC-friendly downstream.

SOURCES SAP HANA Oracle MySQL Kafka CDC ADLS Gen2 REST APIs BRONZE Raw Ingestion append-only · schema on read Autoloader Kafka Structured Streaming JDBC parameterized ingest Checkpoint + offset mgmt Trigger-once / incremental SILVER Cleansed & Conformed MERGE · schema-enforced Deduplication Cleansing · null handling Idempotent MERGE Data quality contracts Source→Target recon GOLD Business-Ready aggregated · SLO'd Z-order / Liquid clustering Spark tuning · AQE off for 8GB+ Aggregations · KPI models Lineage · observability Serving via ADF / SQL OPS ADF Workflows DAB Unity Cat. CI/CD Retry
Bronze · raw
Silver · cleansed
Gold · business
Orchestration
Selected work

Where the pipelines run.

Three enterprise engagements. Different domains, same obsession with idempotency, observability, and not paging the on-call at 3am.

RELIANCE · ENERGY & INFRASTRUCTURE

Enterprise-wide Medallion Lakehouse

Greenfield · 20+ subsidiaries · SAP HANA · Oracle · MySQL · Kafka CDC

First unified analytics platform for a diversified energy & infrastructure group — built a parameterized PySpark ingestion framework that collapsed per-source duplication across the entire Medallion stack.

  • Cut new-source onboarding from ~3 weeks to 3 days via parameterized JDBC ingestion
  • Sub-minute CDC latency via Kafka Structured Streaming with watermark + offset management
  • Release time 2 days → 30 min via Databricks Asset Bundles, DEV→UAT→PROD
  • ~70% drop in data quality incidents via Bronze→Silver transformation contracts
  • Zero-touch ADLS Gen2 ingestion via Autoloader + trigger-once semantics
  • Full lineage observability across 15+ active pipelines via Workflows + structured logging
PySparkDelta LakeAutoloaderKafka CDCDABUnity Catalog
HDFC BANK · BFSI

600-Job Pentaho → PySpark Migration

Air-gapped · RBI-regulated · 80M records · 25 TB/month

Migrated an air-gapped BFSI estate from legacy Pentaho to PySpark on Azure Databricks — delivered the client's first version-controlled CI/CD pipeline and 10× query performance.

  • 10× query perf across 600+ migrated jobs via Spark tuning, partition + Z-order
  • Saved ~240 engineering hours via REST API notebook-cloning automation, run overnight
  • Resolved critical failures on 80M records / 25 TB/month — disabled AQE broadcast for 8GB+ tables
  • ~40% faster release cycle via multi-env CI/CD on Azure DevOps + ARM
  • Met RBI residency via ABFS + folder-level SAS + Managed Delta write buffers
  • Re-architected pipelines exceeding ADF's 40-activity limit into sub-pipelines
PySparkADFAzure DevOpsZ-orderARM TemplatesABFS / SAS
TE CONNECTIVITY · INDUSTRIAL

5-Region Sales Datamart Migration

526M+ records · Oracle PL/SQL → PySpark · zero-defect sign-off

Re-architected five regional sales datamarts into production PySpark — eliminated Control-M dependency, hit latency SLAs, signed off zero-defect against the legacy Oracle output.

  • Processed 526M+ records across Booking, Backlog, and Costed Sales Detail
  • ~35% compute cost reduction via native Spark replacements + transform consolidation
  • Eliminated Control-M scheduler dependency via Python-native Databricks orchestration
  • Zero-defect production sign-off via row-count + 40+ business-metric reconciliation
  • Replicated full Linux shell conditional branching in Python without losing semantics
  • Native Spark UDFs over custom ones — fewer serialization hot spots
PySparkDelta LakeSpark TuningValidation FrameworkWorkflows
Career

Where I've shipped.

Senior Data Engineer — Celebal Technologies, Pune

Jan 2022 — Present · 4+ years

Engagements across Reliance (Energy & Infrastructure), HDFC Bank (BFSI), and TE Connectivity (Industrial Sales). Full pipeline lifecycle ownership — Autoloader ingestion, PySpark transformation, Delta Lake optimization, ADF orchestration, CI/CD via Databricks Asset Bundles.

B.Tech, Computer Science & Engineering

Centurion University · CGPA 9.2/10 · Apr 2016

Foundations in CS that I'd later mostly use to debug Spark serialization issues.

PG Diploma, Big Data Analytics

C-DAC, Pune · Sep 2021

Where the PySpark obsession started. Hadoop was already dying. Spark + Delta was the obvious next thing.

Certifications

Databricks Certified Data Engineer Professional
Databricks · issued
Databricks Certified Data Engineer Associate
Databricks · issued

Education

B.Tech, Computer Science & Engineering
Centurion University · CGPA 9.2/10 · Apr 2016
PG Diploma, Big Data Analytics
C-DAC, Pune · Sep 2021
Let's talk pipelines

Hiring? Building a lakehouse? Paging me?

Open to Senior Data Engineer and Platform Data roles. Open to contract work, hard problems, and teams that ship.