
🛠️ DevOps & SRE · State Machine
The stages of a production incident from the first alert to the postmortem, with severity escalation along the way.
Drawing diagram…
State machine for an incident: Triggered by an alert, Acknowledged by on-call, Investigating, Escalated if severity rises, Mitigated when users are no longer affected, Resolved when the root cause is fixed, then PostmortemDone. False alarms are closed right after acknowledgement.
stateDiagram-v2 [*] --> Triggered: Alert fires Triggered --> Acknowledged: On-call responds Triggered --> Escalated: No ack in 10 minutes Acknowledged --> Investigating Acknowledged --> Closed: False alarm Investigating --> Escalated: Severity raised Escalated --> Investigating: More responders join Investigating --> Mitigated: Users no longer affected Mitigated --> Resolved: Root cause fixed Resolved --> PostmortemDone: Review published PostmortemDone --> [*] Closed --> [*]
How a code change reaches production with GitOps: CI builds and tests, the image tag is written to a config repo, and Argo CD syncs the cluster to match Git.
A complete observability setup: OpenTelemetry collects metrics, logs and traces from services, which are stored separately and viewed together in Grafana with alerting.
Who does what when production breaks: monitoring pages the on-call engineer, an incident channel is opened, customers are updated and the fix is rolled out.
A blue-green setup where a new version is deployed next to the live one and traffic is switched only after it passes checks, allowing instant rollback.
How infrastructure changes are made safely with Terraform: plan on pull request, policy checks, approval, apply and drift detection.
A self-service platform for developers: a Backstage portal with templates that create repos, pipelines and environments without tickets to the ops team.