
🛠️ DevOps & SRE · C4 Architecture · Pro
A complete observability setup: OpenTelemetry collects metrics, logs and traces from services, which are stored separately and viewed together in Grafana with alerting.
C4 Architecture diagrams are part of Pro. Anyone can view this one; generating and editing it needs Pro.
Drawing diagram…
Observability platform: services are instrumented with OpenTelemetry SDKs and send data to an OpenTelemetry Collector. Metrics go to Prometheus (with Mimir for long-term storage), logs to Loki and traces to Tempo. Grafana shows dashboards that link traces to logs. Alertmanager sends alerts to PagerDuty and Slack.
C4Container
title Observability Stack
Person(sre, "SRE / Developer", "Investigates issues")
System_Boundary(obs, "Observability Platform") {
Container(col, "OTel Collector", "OpenTelemetry", "Receives all signals")
ContainerDb(prom, "Metrics", "Prometheus + Mimir", "Time series")
ContainerDb(loki, "Logs", "Loki", "Log streams")
ContainerDb(tempo, "Traces", "Tempo", "Distributed traces")
Container(graf, "Dashboards", "Grafana", "Metrics, logs, traces together")
Container(am, "Alerting", "Alertmanager", "Routes alerts")
}
System_Ext(svc, "Microservices", "Instrumented with OTel SDK")
System_Ext(pd, "PagerDuty", "On-call paging")
Rel(svc, col, "Metrics, logs, traces", "OTLP")
Rel(col, prom, "Metrics")
Rel(col, loki, "Logs")
Rel(col, tempo, "Traces")
Rel(graf, prom, "Queries")
Rel(graf, loki, "Queries")
Rel(graf, tempo, "Queries")
Rel(prom, am, "Firing alerts")
Rel(am, pd, "Pages on-call")
Rel(sre, graf, "Uses")How a code change reaches production with GitOps: CI builds and tests, the image tag is written to a config repo, and Argo CD syncs the cluster to match Git.
The stages of a production incident from the first alert to the postmortem, with severity escalation along the way.
Who does what when production breaks: monitoring pages the on-call engineer, an incident channel is opened, customers are updated and the fix is rolled out.
A blue-green setup where a new version is deployed next to the live one and traffic is switched only after it passes checks, allowing instant rollback.
How infrastructure changes are made safely with Terraform: plan on pull request, policy checks, approval, apply and drift detection.
A self-service platform for developers: a Backstage portal with templates that create repos, pipelines and environments without tickets to the ops team.