CloudSketch AI Logo
FeaturesTemplatesPricingEnterpriseAboutContactLog in
πŸŒ™β˜€οΈ
Log inStart free
Home/Templates/Data Engineering & Analytics

πŸ“Š Data Engineering & Analytics Β· Sankey

Daily Data Volume Flow

Where a company's daily data comes from and where it ends up, measured in gigabytes, to plan storage and processing costs.

More Data Engineering & Analytics templates

Drawing diagram…

What this diagram shows

  • Events are the largest source of data
  • Most raw data is archived, not queried
  • Only a small share reaches dashboards

Prompt used

Sankey of daily data volume in GB: app events 400, database changes 150, logs 300 and partner files 50 go into ingestion. From ingestion, 500 GB goes to the raw archive, 300 to the cleaned lake and 100 to real-time streams. From the cleaned lake, 120 GB becomes warehouse tables and 180 feeds ML features. Warehouse tables feed 40 GB of dashboard aggregates.

Mermaid code
sankey-beta
App events,Ingestion,400
Database changes,Ingestion,150
Logs,Ingestion,300
Partner files,Ingestion,50
Ingestion,Raw archive,500
Ingestion,Clean lake,300
Ingestion,Real-time streams,100
Clean lake,Warehouse tables,120
Clean lake,ML features,180
Warehouse tables,Dashboard aggregates,40
Warehouse tables,Ad-hoc queries,80

Related templates

C4 ArchitectureProData Engineering & Analytics

Modern Data Platform Architecture

A typical modern data stack: data is loaded from apps and SaaS tools into a cloud warehouse, modelled with dbt, orchestrated with Airflow and served to BI dashboards.

Data FlowData Engineering & Analytics

Change Data Capture Pipeline

How every insert, update and delete in an operational database is streamed to the data lake and search index in near real time using change data capture.

Database ERDData Engineering & Analytics

Sales Star Schema

A classic star schema for sales analytics: one fact table of order lines surrounded by date, customer, product, store and promotion dimensions.

FlowchartData Engineering & Analytics

Nightly Batch ETL Workflow

The steps of a nightly batch pipeline with data quality gates: extract, validate, transform, load and publish, stopping safely when checks fail.

DeploymentProData Engineering & Analytics

Real-time Streaming Analytics Deployment

Where a real-time analytics stack runs: Kafka for events, Flink for stream processing, a real-time OLAP database and live dashboards.

State MachineData Engineering & Analytics

Data Pipeline Run States

The states of a single pipeline run in an orchestrator like Airflow, including retries, upstream failures and manual reruns.