
π Data Engineering & Analytics Β· Data Flow
The medallion pattern for a data lake: raw data lands in bronze, is cleaned in silver, and turned into business-ready tables in gold.
Drawing diagramβ¦
Medallion architecture data flow: batch files, API data and streaming events land unchanged in the bronze zone of a Delta Lake. A cleaning job de-duplicates, standardises types and masks personal data into silver. Aggregation jobs build gold tables such as daily revenue and customer 360. BI tools and ML features read from gold.
flowchart LR F[Batch Files] -->|As received| BR[(Bronze Zone)] API[Partner APIs] -->|JSON| BR EV[Event Stream] -->|Events| BR BR -->|Raw records| CL[Cleaning Job] CL -->|De-duplicated, typed, masked| SV[(Silver Zone)] SV -->|Clean tables| AG[Aggregation Jobs] AG -->|Daily revenue, customer 360| GD[(Gold Zone)] GD -->|Metrics| BI[BI Dashboards] GD -->|Features| ML[ML Training]
A typical modern data stack: data is loaded from apps and SaaS tools into a cloud warehouse, modelled with dbt, orchestrated with Airflow and served to BI dashboards.
How every insert, update and delete in an operational database is streamed to the data lake and search index in near real time using change data capture.
A classic star schema for sales analytics: one fact table of order lines surrounded by date, customer, product, store and promotion dimensions.
The steps of a nightly batch pipeline with data quality gates: extract, validate, transform, load and publish, stopping safely when checks fail.
Where a real-time analytics stack runs: Kafka for events, Flink for stream processing, a real-time OLAP database and live dashboards.
The states of a single pipeline run in an orchestrator like Airflow, including retries, upstream failures and manual reruns.