Databricks framework to validate Data Quality of pySpark DataFrames and Tables
-
Updated
Sep 9, 2026 - Python
Databricks framework to validate Data Quality of pySpark DataFrames and Tables
Metadata-driven framework for Databricks Spark Declarative Pipelines. Config-driven, pattern based approach to batch & streaming across the medallion architecture. Deploys via Declarative Automation Bundles. Built for simplicity, extensibility, and alignment with the Databricks product roadmap.
Open-source community study guide for all six Databricks certifications (Data Engineer Associate / Professional, Data Analyst Associate, ML Associate / Professional, GenAI Engineer Associate). Aligned to the 2025-2026 official exam guides. Obsidian-flavoured Markdown; PRs welcome.
Medallion Architecture for Data Engineering projects
Medallion analytics for the Ceres Open Data Index on Databricks — Lakeflow Declarative Pipelines (Bronze → Silver → Gold), runs on Free Edition.
Databricks-native data trust pipeline — intake certification, drift gating, and control benchmarking in a single deployable product.
Bring your Claude Code skill unchanged and run it as a governed Databricks job. Publish once to a Unity Catalog volume, reuse from any job, chain skills into a pipeline (markdown in, branded PowerPoint out). No external API key. Runs on Free Edition.
Databricks SQL in Action — End-to-end medallion architecture lab using Unity Catalog, Volumes, Streaming Tables, Materialized Views, AI SQL functions, dashboards, lineage, and workflow orchestration.
Built an end-to-end retail lakehouse using the Medallion architecture with batch and streaming pipelines for scalable business analytics.
Hands-on Azure Databricks learning project — medallion pipeline, Lakeflow SDP, Delta Lake, Unity Catalog security, SQL analytics and Jobs. Built end-to-end on Azure with real executed notebooks.
Demo of Databricks Lakeflow Jobs Automation with StackQL and Databricks Asset Bundles
Metadata-driven schemaless MongoDB Debezium CDC ingestion into a Unity Catalog medallion lakehouse on Databricks Lakeflow Declarative Pipelines, using the VARIANT data type
Sample Databricks Asset Bundle: hotel daily performance KPIs with Lakeflow SDP, UC Metric Views, AI/BI dashboard (Brickstar styled), and Genie NL→SQL
Built a metadata-driven Azure lakehouse that incrementally ingests Spotify-style SQL data with ADF, processes new Parquet files with Databricks Auto Loader, and publishes SCD-managed Delta facts and dimensions.
Built a real-time Azure lakehouse that streams synthetic ride-booking events from FastAPI through Event Hubs into Databricks, unifies them with historical data, and publishes SCD-managed facts and dimensions.
Tsuga Logs community connector for Databricks Lakeflow Connect — scheduled log ingestion into Delta tables you own
End-to-end Azure Databricks Data Engineering Pipeline with Medallion Architecture, Delta Lake, Unity Catalog, and Lakeflow Jobs.
Event-driven lakehouse on Databricks with real CDC from PostgreSQL via Debezium — medallion architecture, Lakeflow Declarative Pipelines, SCD, CDF, and liquid clustering. Infra as code with Terraform.
Built an OAuth-authenticated Spark streaming lakehouse that consumes NASA GCN Fermi gamma-ray burst notices from Kafka, parses classic-text messages, and publishes a governed analytical snowflake schema.
To associate your repository with the lakeflow topic, visit your repo's landing page and select "manage topics."