
🛠️ DevOps & SRE · Flowchart
What to do when a release goes bad: decide quickly between a feature-flag switch-off, a rollback and a hotfix.
Drawing diagram…
Deployment rollback runbook flowchart: after a release the error rate or latency rises. Check if the release is the cause by comparing with the deploy time. If not, follow the normal incident flow. If yes and the change is behind a feature flag, switch the flag off. Otherwise check if the release ran a database migration that is not backward compatible: if not, roll back to the previous version in the CD tool; if yes, ship a hotfix forward. Verify metrics return to normal, freeze further deploys, and open a ticket to fix and add a test.
flowchart TD
A(["Error rate or latency up after release"]) --> B{"Started at deploy time?"}
B -->|"No"| C["Follow incident response flow"]
B -->|"Yes"| D{"Change behind a feature flag?"}
D -->|"Yes"| E["Switch the flag off"]
D -->|"No"| F{"Non-backward-compatible<br/>DB migration?"}
F -->|"No"| G["Roll back to previous version<br/>in the CD tool"]
F -->|"Yes"| H["Ship a hotfix forward"]
E --> I{"Metrics back to normal?"}
G --> I
H --> I
I -->|"No"| C
I -->|"Yes"| J["Freeze deploys for this service"]
J --> K(["Ticket: fix and add a test"])How a code change reaches production with GitOps: CI builds and tests, the image tag is written to a config repo, and Argo CD syncs the cluster to match Git.
A complete observability setup: OpenTelemetry collects metrics, logs and traces from services, which are stored separately and viewed together in Grafana with alerting.
The stages of a production incident from the first alert to the postmortem, with severity escalation along the way.
Who does what when production breaks: monitoring pages the on-call engineer, an incident channel is opened, customers are updated and the fix is rolled out.
A blue-green setup where a new version is deployed next to the live one and traffic is switched only after it passes checks, allowing instant rollback.
How infrastructure changes are made safely with Terraform: plan on pull request, policy checks, approval, apply and drift detection.