
π Data Engineering & Analytics Β· Sankey
Where a company's daily data comes from and where it ends up, measured in gigabytes, to plan storage and processing costs.
Drawing diagramβ¦
Sankey of daily data volume in GB: app events 400, database changes 150, logs 300 and partner files 50 go into ingestion. From ingestion, 500 GB goes to the raw archive, 300 to the cleaned lake and 100 to real-time streams. From the cleaned lake, 120 GB becomes warehouse tables and 180 feeds ML features. Warehouse tables feed 40 GB of dashboard aggregates.
sankey-beta App events,Ingestion,400 Database changes,Ingestion,150 Logs,Ingestion,300 Partner files,Ingestion,50 Ingestion,Raw archive,500 Ingestion,Clean lake,300 Ingestion,Real-time streams,100 Clean lake,Warehouse tables,120 Clean lake,ML features,180 Warehouse tables,Dashboard aggregates,40 Warehouse tables,Ad-hoc queries,80
A typical modern data stack: data is loaded from apps and SaaS tools into a cloud warehouse, modelled with dbt, orchestrated with Airflow and served to BI dashboards.
How every insert, update and delete in an operational database is streamed to the data lake and search index in near real time using change data capture.
A classic star schema for sales analytics: one fact table of order lines surrounded by date, customer, product, store and promotion dimensions.
The steps of a nightly batch pipeline with data quality gates: extract, validate, transform, load and publish, stopping safely when checks fail.
Where a real-time analytics stack runs: Kafka for events, Flink for stream processing, a real-time OLAP database and live dashboards.
The states of a single pipeline run in an orchestrator like Airflow, including retries, upstream failures and manual reruns.