
π Data Engineering & Analytics Β· State Machine
The states of a single pipeline run in an orchestrator like Airflow, including retries, upstream failures and manual reruns.
Drawing diagramβ¦
State machine for a pipeline run: Scheduled, Queued, Running, then Success. On error go to UpForRetry and back to Running up to 3 times, then Failed. If an upstream dependency failed the run is UpstreamFailed. Engineers can clear Failed or Success runs to rerun them.
stateDiagram-v2 [*] --> Scheduled Scheduled --> Queued: Start time reached Scheduled --> UpstreamFailed: Dependency failed Queued --> Running: Worker free Running --> Success Running --> UpForRetry: Task error UpForRetry --> Running: Retry, max 3 UpForRetry --> Failed: Retries used up Failed --> Queued: Cleared by engineer Success --> Queued: Rerun requested Success --> [*] UpstreamFailed --> [*]
A typical modern data stack: data is loaded from apps and SaaS tools into a cloud warehouse, modelled with dbt, orchestrated with Airflow and served to BI dashboards.
How every insert, update and delete in an operational database is streamed to the data lake and search index in near real time using change data capture.
A classic star schema for sales analytics: one fact table of order lines surrounded by date, customer, product, store and promotion dimensions.
The steps of a nightly batch pipeline with data quality gates: extract, validate, transform, load and publish, stopping safely when checks fail.
Where a real-time analytics stack runs: Kafka for events, Flink for stream processing, a real-time OLAP database and live dashboards.
The medallion pattern for a data lake: raw data lands in bronze, is cleaned in silver, and turned into business-ready tables in gold.