Skip to content
Back to projects
Databricks dashboard showing churn rate, model accuracy, calibration and performance by segment
Big Data · MLOpsApplied academic project2026

Telco churn Lakehouse and MLOps

Co-developed end-to-end data platform that predicts customer churn for a telecom operator on synthetic data using PySpark, Delta Lake, Unity Catalog and MLflow.

  • DatabricksLakehouse
  • Asset Bundlesdatabricks.yml
  • PipelinesMedallion serverless
  • Delta LakeBronze · Silver · Gold
  • SparkMLlib
  • MLflowTracking + Registry
  • Unity CatalogModels + Tables
  • LicenseMIT

Problem

Build a reproducible data and ML lifecycle for predicting customer churn while making the synthetic nature of the dataset explicit.

Architecture

Databricks Lakehouse with incremental ingestion, a Delta Lake Medallion architecture, temporal feature engineering, MLflow, Unity Catalog and monitored batch inference.

Data flow

Synthetic data → Auto Loader → Bronze/Silver/Gold → point-in-time dataset → training and registry → simulated inference and monitoring.

Results

Contribution and authorship

  • End-to-end co-implementation confirmed by Alonso on 2026-09-14
  • MLflow experiment, three-task ML job and simulation run associated with Alonso
  • Shared work across the Medallion pipeline, model lifecycle and daily validation
Alonso Marcos Muñoz · End-to-end co-authorJose Barros · Co-author

Evidence

Executed Medallion DAG
Executed Medallion DAG
Lakehouse monitor
Lakehouse monitor

Limitations

  • • The data is synthetic and the production-labelled scenario is an academic simulation
  • • Authenticated validation and deployment require an active Databricks workspace
  • • Authorship is shared and both co-authors remain credited

Stack

https://github.com/alonsomarcosm99/databricks-telco-churn-lakehouse