CloudSketch AI Logo
FeaturesTemplatesPricingEnterpriseAboutContactLog in
🌙☀️
Log inStart free
Home/Templates/DevOps & SRE

🛠️ DevOps & SRE · Flowchart

Database Failover Runbook

Step-by-step runbook for failing over from a broken primary database to a replica, with checks before and after.

More DevOps & SRE templates

Drawing diagram…

What this diagram shows

  • Confirm the primary is really down before failing over
  • Check replica lag to know how much data could be lost
  • Freeze writes, promote, repoint, verify, then rebuild the old primary

Prompt used

Database failover runbook flowchart: alert that the primary database is unreachable. Confirm from two places (monitoring and a direct connection). If it is a network blip, wait and recheck. If down, check replica lag: under 5 seconds proceed, otherwise get approval from the incident commander for possible data loss. Put the app in read-only mode, promote the replica, update the connection string or DNS, restart app pods, run smoke tests. If smoke tests fail, roll back to read-only and escalate to the DBA. If they pass, end read-only mode, rebuild the old primary as a new replica, and record the timeline.

Mermaid code
flowchart TD
  A(["Alert: primary DB unreachable"]) --> B["Confirm from monitoring<br/>and a direct connection"]
  B --> C{"Really down?"}
  C -->|"No, network blip"| D["Wait 2 min and recheck"]
  D --> B
  C -->|"Yes"| E{"Replica lag under 5 s?"}
  E -->|"No"| F["Incident commander approves<br/>possible data loss"]
  E -->|"Yes"| G["Switch app to read-only"]
  F --> G
  G --> H["Promote replica to primary"]
  H --> I["Update connection string / DNS"]
  I --> J["Restart app pods"]
  J --> K{"Smoke tests pass?"}
  K -->|"No"| L["Stay read-only<br/>escalate to DBA"]
  K -->|"Yes"| M["End read-only mode"]
  M --> N["Rebuild old primary<br/>as new replica"]
  N --> O(["Record timeline for postmortem"])

Related templates

FlowchartDevOps & SRE

GitOps Deployment Flow

How a code change reaches production with GitOps: CI builds and tests, the image tag is written to a config repo, and Argo CD syncs the cluster to match Git.

C4 ArchitectureProDevOps & SRE

Observability Stack (Metrics, Logs, Traces)

A complete observability setup: OpenTelemetry collects metrics, logs and traces from services, which are stored separately and viewed together in Grafana with alerting.

State MachineDevOps & SRE

Incident Lifecycle

The stages of a production incident from the first alert to the postmortem, with severity escalation along the way.

SequenceDevOps & SRE

Incident Response Sequence

Who does what when production breaks: monitoring pages the on-call engineer, an incident channel is opened, customers are updated and the fix is rolled out.

DeploymentProDevOps & SRE

Blue-Green Deployment

A blue-green setup where a new version is deployed next to the live one and traffic is switched only after it passes checks, allowing instant rollback.

FlowchartDevOps & SRE

Infrastructure as Code Pipeline

How infrastructure changes are made safely with Terraform: plan on pull request, policy checks, approval, apply and drift detection.