
📊 Data Engineering & Analytics · Deployment · Pro
Where a real-time analytics stack runs: Kafka for events, Flink for stream processing, a real-time OLAP database and live dashboards.
Deployment diagrams are part of Pro. Anyone can view this one; generating and editing it needs Pro.
Drawing diagram…
Streaming analytics on AWS: apps send clickstream events to Amazon MSK (Kafka) across three AZs. Flink on Kubernetes (EKS) aggregates events in 1-minute windows with checkpoints in S3. Results go to Apache Pinot for real-time queries and to S3 for history. Grafana and a React dashboard show live metrics.
flowchart TB
APPS[Web and Mobile Apps] --> MSK
subgraph AWS[AWS Region eu-west-1]
subgraph MSK[Amazon MSK - 3 AZs]
B1[Broker AZ-a]
B2[Broker AZ-b]
B3[Broker AZ-c]
end
subgraph EKS[EKS Cluster]
FJM[Flink JobManager]
FTM[Flink TaskManagers x6]
PIN[Apache Pinot]
DASH[Live Dashboard]
GRAF[Grafana]
end
S3[(S3 - checkpoints and history)]
end
MSK --> FTM
FJM --> FTM
FTM --> S3
FTM --> PIN
DASH --> PIN
GRAF --> PINA typical modern data stack: data is loaded from apps and SaaS tools into a cloud warehouse, modelled with dbt, orchestrated with Airflow and served to BI dashboards.
How every insert, update and delete in an operational database is streamed to the data lake and search index in near real time using change data capture.
A classic star schema for sales analytics: one fact table of order lines surrounded by date, customer, product, store and promotion dimensions.
The steps of a nightly batch pipeline with data quality gates: extract, validate, transform, load and publish, stopping safely when checks fail.
The states of a single pipeline run in an orchestrator like Airflow, including retries, upstream failures and manual reruns.
The medallion pattern for a data lake: raw data lands in bronze, is cleaned in silver, and turned into business-ready tables in gold.