
🛠️ DevOps & SRE · Flowchart
Step-by-step runbook for failing over from a broken primary database to a replica, with checks before and after.
Drawing diagram…
Database failover runbook flowchart: alert that the primary database is unreachable. Confirm from two places (monitoring and a direct connection). If it is a network blip, wait and recheck. If down, check replica lag: under 5 seconds proceed, otherwise get approval from the incident commander for possible data loss. Put the app in read-only mode, promote the replica, update the connection string or DNS, restart app pods, run smoke tests. If smoke tests fail, roll back to read-only and escalate to the DBA. If they pass, end read-only mode, rebuild the old primary as a new replica, and record the timeline.
flowchart TD
A(["Alert: primary DB unreachable"]) --> B["Confirm from monitoring<br/>and a direct connection"]
B --> C{"Really down?"}
C -->|"No, network blip"| D["Wait 2 min and recheck"]
D --> B
C -->|"Yes"| E{"Replica lag under 5 s?"}
E -->|"No"| F["Incident commander approves<br/>possible data loss"]
E -->|"Yes"| G["Switch app to read-only"]
F --> G
G --> H["Promote replica to primary"]
H --> I["Update connection string / DNS"]
I --> J["Restart app pods"]
J --> K{"Smoke tests pass?"}
K -->|"No"| L["Stay read-only<br/>escalate to DBA"]
K -->|"Yes"| M["End read-only mode"]
M --> N["Rebuild old primary<br/>as new replica"]
N --> O(["Record timeline for postmortem"])How a code change reaches production with GitOps: CI builds and tests, the image tag is written to a config repo, and Argo CD syncs the cluster to match Git.
A complete observability setup: OpenTelemetry collects metrics, logs and traces from services, which are stored separately and viewed together in Grafana with alerting.
The stages of a production incident from the first alert to the postmortem, with severity escalation along the way.
Who does what when production breaks: monitoring pages the on-call engineer, an incident channel is opened, customers are updated and the fix is rolled out.
A blue-green setup where a new version is deployed next to the live one and traffic is switched only after it passes checks, allowing instant rollback.
How infrastructure changes are made safely with Terraform: plan on pull request, policy checks, approval, apply and drift detection.