
π Data Engineering & Analytics Β· Flowchart
The steps of a nightly batch pipeline with data quality gates: extract, validate, transform, load and publish, stopping safely when checks fail.
Drawing diagramβ¦
Nightly ETL flowchart: at 1 am extract files from SFTP and the ERP, check row counts and schemas, and if checks fail alert on-call and stop. Otherwise transform in Spark, run data quality tests (nulls, duplicates, totals match source). If tests pass, load into staging tables, swap them into production and refresh dashboards; if not, keep yesterday's tables and alert.
flowchart TD
A[1 am schedule starts] --> B[Extract ERP and SFTP files]
B --> C{Row counts and schema OK?}
C -->|No| X[Alert on-call and stop]
C -->|Yes| D[Transform in Spark]
D --> E[Run quality tests]
E --> F{Nulls, duplicates and totals OK?}
F -->|No| Y[Keep yesterday's tables and alert]
F -->|Yes| G[Load staging tables]
G --> H[Swap staging into production]
H --> I[Refresh dashboards]
I --> J[Post run summary to Slack]A typical modern data stack: data is loaded from apps and SaaS tools into a cloud warehouse, modelled with dbt, orchestrated with Airflow and served to BI dashboards.
How every insert, update and delete in an operational database is streamed to the data lake and search index in near real time using change data capture.
A classic star schema for sales analytics: one fact table of order lines surrounded by date, customer, product, store and promotion dimensions.
Where a real-time analytics stack runs: Kafka for events, Flink for stream processing, a real-time OLAP database and live dashboards.
The states of a single pipeline run in an orchestrator like Airflow, including retries, upstream failures and manual reruns.
The medallion pattern for a data lake: raw data lands in bronze, is cleaned in silver, and turned into business-ready tables in gold.