
🛠️ DevOps & SRE · Flowchart
From the first alert to the postmortem: triage by severity, assign an incident commander, mitigate, communicate and learn.
Drawing diagram…
Incident response flowchart: an alert fires or a customer reports a problem. The on-call engineer acknowledges within 5 minutes and assesses severity (SEV1 full outage, SEV2 major feature down, SEV3 minor). SEV1/SEV2: page the incident commander, open an incident channel and a status page update, post updates every 30 minutes. SEV3: create a ticket and fix in working hours. Mitigate (roll back, fail over, scale up or disable the feature flag); if not mitigated in 30 minutes escalate to the service owner. Once resolved, monitor for 30 minutes, close the status page, and for SEV1/SEV2 hold a blameless postmortem within 5 days with action items tracked.
flowchart TD
A(["Alert fires or customer report"]) --> B["On-call acknowledges<br/>within 5 min"]
B --> C{"Severity?"}
C -->|"SEV1 outage / SEV2 major"| D["Page incident commander"]
C -->|"SEV3 minor"| E["Create ticket<br/>fix in working hours"]
D --> F["Open incident channel<br/>and status page"]
F --> G["Mitigate: roll back, fail over,<br/>scale up or disable flag"]
G --> H{"Mitigated within 30 min?"}
H -->|"No"| I["Escalate to service owner"]
I --> G
H -->|"Yes"| J["Monitor for 30 min"]
F -.->|"every 30 min"| K["Post status update"]
J --> L["Resolve and close status page"]
L --> M["Blameless postmortem<br/>within 5 days"]
M --> N(["Action items tracked to done"])
E --> O(["Closed"])How a code change reaches production with GitOps: CI builds and tests, the image tag is written to a config repo, and Argo CD syncs the cluster to match Git.
A complete observability setup: OpenTelemetry collects metrics, logs and traces from services, which are stored separately and viewed together in Grafana with alerting.
The stages of a production incident from the first alert to the postmortem, with severity escalation along the way.
Who does what when production breaks: monitoring pages the on-call engineer, an incident channel is opened, customers are updated and the fix is rolled out.
A blue-green setup where a new version is deployed next to the live one and traffic is switched only after it passes checks, allowing instant rollback.
How infrastructure changes are made safely with Terraform: plan on pull request, policy checks, approval, apply and drift detection.