CloudSketch AI Logo
FeaturesTemplatesPricingEnterpriseAboutContactLog in
🌙☀️
Log inStart free
Home/Templates/DevOps & SRE

🛠️ DevOps & SRE · Flowchart

Disk Full Alert Runbook

Runbook for a disk-usage alert on a server or volume: find what is growing, free space safely, then stop it happening again.

More DevOps & SRE templates

Drawing diagram…

What this diagram shows

  • Find the cause before deleting anything
  • Safe clean-ups first: old logs, temp files, old images
  • Expand the volume only when growth is real

Prompt used

Disk full alert runbook flowchart: alert that disk usage is above 90%. Find the largest directories. If it is logs, rotate and compress them and check log retention. If it is temp files or old container images, clean them up. If it is real data growth, expand the volume and plan capacity. If a process holds deleted files, restart it. Check usage is under 70%; if not, escalate to the service owner. Finally add an alert at 80% and a ticket for the root cause.

Mermaid code
flowchart TD
  A(["Alert: disk usage above 90%"]) --> B["Find largest directories"]
  B --> C{"What is growing?"}
  C -->|"Logs"| D["Rotate and compress logs<br/>check retention"]
  C -->|"Temp files / old images"| E["Clean temp files<br/>and old container images"]
  C -->|"Real data growth"| F["Expand the volume<br/>plan capacity"]
  C -->|"Deleted files still open"| G["Restart the process holding them"]
  D --> H{"Usage under 70%?"}
  E --> H
  F --> H
  G --> H
  H -->|"No"| I["Escalate to service owner"]
  H -->|"Yes"| J["Add an 80% early alert"]
  J --> K(["Ticket for root cause"])

Related templates

FlowchartDevOps & SRE

GitOps Deployment Flow

How a code change reaches production with GitOps: CI builds and tests, the image tag is written to a config repo, and Argo CD syncs the cluster to match Git.

C4 ArchitectureProDevOps & SRE

Observability Stack (Metrics, Logs, Traces)

A complete observability setup: OpenTelemetry collects metrics, logs and traces from services, which are stored separately and viewed together in Grafana with alerting.

State MachineDevOps & SRE

Incident Lifecycle

The stages of a production incident from the first alert to the postmortem, with severity escalation along the way.

SequenceDevOps & SRE

Incident Response Sequence

Who does what when production breaks: monitoring pages the on-call engineer, an incident channel is opened, customers are updated and the fix is rolled out.

DeploymentProDevOps & SRE

Blue-Green Deployment

A blue-green setup where a new version is deployed next to the live one and traffic is switched only after it passes checks, allowing instant rollback.

FlowchartDevOps & SRE

Infrastructure as Code Pipeline

How infrastructure changes are made safely with Terraform: plan on pull request, policy checks, approval, apply and drift detection.