Enterprise-wide Medallion Lakehouse
First unified analytics platform for a diversified energy & infrastructure group — built a parameterized PySpark ingestion framework that collapsed per-source duplication across the entire Medallion stack.
- Cut new-source onboarding from ~3 weeks to 3 days via parameterized JDBC ingestion
- Sub-minute CDC latency via Kafka Structured Streaming with watermark + offset management
- Release time 2 days → 30 min via Databricks Asset Bundles, DEV→UAT→PROD
- ~70% drop in data quality incidents via Bronze→Silver transformation contracts
- Zero-touch ADLS Gen2 ingestion via Autoloader + trigger-once semantics
- Full lineage observability across 15+ active pipelines via Workflows + structured logging