CloudSketch AI Logo
FeaturesTemplatesPricingEnterpriseAboutContactLog in
🌙☀️
Log inStart free
Home/Templates/DevOps & SRE

🛠️ DevOps & SRE · Flowchart

Deployment Rollback Runbook

What to do when a release goes bad: decide quickly between a feature-flag switch-off, a rollback and a hotfix.

More DevOps & SRE templates

Drawing diagram…

What this diagram shows

  • Feature flag first: fastest and safest
  • Database migrations decide whether a rollback is safe
  • Hotfix only when rollback is not possible

Prompt used

Deployment rollback runbook flowchart: after a release the error rate or latency rises. Check if the release is the cause by comparing with the deploy time. If not, follow the normal incident flow. If yes and the change is behind a feature flag, switch the flag off. Otherwise check if the release ran a database migration that is not backward compatible: if not, roll back to the previous version in the CD tool; if yes, ship a hotfix forward. Verify metrics return to normal, freeze further deploys, and open a ticket to fix and add a test.

Mermaid code
flowchart TD
  A(["Error rate or latency up after release"]) --> B{"Started at deploy time?"}
  B -->|"No"| C["Follow incident response flow"]
  B -->|"Yes"| D{"Change behind a feature flag?"}
  D -->|"Yes"| E["Switch the flag off"]
  D -->|"No"| F{"Non-backward-compatible<br/>DB migration?"}
  F -->|"No"| G["Roll back to previous version<br/>in the CD tool"]
  F -->|"Yes"| H["Ship a hotfix forward"]
  E --> I{"Metrics back to normal?"}
  G --> I
  H --> I
  I -->|"No"| C
  I -->|"Yes"| J["Freeze deploys for this service"]
  J --> K(["Ticket: fix and add a test"])

Related templates

FlowchartDevOps & SRE

GitOps Deployment Flow

How a code change reaches production with GitOps: CI builds and tests, the image tag is written to a config repo, and Argo CD syncs the cluster to match Git.

C4 ArchitectureProDevOps & SRE

Observability Stack (Metrics, Logs, Traces)

A complete observability setup: OpenTelemetry collects metrics, logs and traces from services, which are stored separately and viewed together in Grafana with alerting.

State MachineDevOps & SRE

Incident Lifecycle

The stages of a production incident from the first alert to the postmortem, with severity escalation along the way.

SequenceDevOps & SRE

Incident Response Sequence

Who does what when production breaks: monitoring pages the on-call engineer, an incident channel is opened, customers are updated and the fix is rolled out.

DeploymentProDevOps & SRE

Blue-Green Deployment

A blue-green setup where a new version is deployed next to the live one and traffic is switched only after it passes checks, allowing instant rollback.

FlowchartDevOps & SRE

Infrastructure as Code Pipeline

How infrastructure changes are made safely with Terraform: plan on pull request, policy checks, approval, apply and drift detection.