
π Data Engineering & Analytics Β· Sequence
How a pipeline checks data quality before publishing: expectations run against the new data, results are stored, and bad data is quarantined.
Drawing diagramβ¦
Sequence for data quality: after loading, Airflow asks Great Expectations to validate the new orders batch. Great Expectations reads the batch from the warehouse and checks rules such as no null customer IDs and totals within 5 percent of source. Results are saved. If all pass Airflow publishes the table. If some fail, bad rows are moved to a quarantine table and the data owner gets a Slack alert.
sequenceDiagram
participant AF as Airflow
participant GE as Great Expectations
participant WH as Warehouse
participant RES as Results Store
participant SL as Slack
AF->>GE: Validate new orders batch
GE->>WH: Read batch
WH-->>GE: Rows
GE->>GE: No null customer IDs, totals within 5%
GE->>RES: Save results
alt All checks pass
GE-->>AF: Passed
AF->>WH: Publish to reporting schema
else Some checks fail
GE-->>AF: Failed checks
AF->>WH: Move bad rows to quarantine
AF->>SL: Alert data owner with report link
endA typical modern data stack: data is loaded from apps and SaaS tools into a cloud warehouse, modelled with dbt, orchestrated with Airflow and served to BI dashboards.
How every insert, update and delete in an operational database is streamed to the data lake and search index in near real time using change data capture.
A classic star schema for sales analytics: one fact table of order lines surrounded by date, customer, product, store and promotion dimensions.
The steps of a nightly batch pipeline with data quality gates: extract, validate, transform, load and publish, stopping safely when checks fail.
Where a real-time analytics stack runs: Kafka for events, Flink for stream processing, a real-time OLAP database and live dashboards.
The states of a single pipeline run in an orchestrator like Airflow, including retries, upstream failures and manual reruns.