CloudSketch AI Logo
FeaturesTemplatesPricingEnterpriseAboutContactLog in
🌙☀️
Log inStart free
Home/Templates/DevOps & SRE

🛠️ DevOps & SRE

DevOps & SRE architecture diagram templates

GitOps, observability, incident response, on-call runbooks, release strategies and infrastructure as code. Open any template to see the finished diagram, then use it as the starting point for your own in the Studio.

All industries🏥 Healthcare🏦 Finance & Banking🛒 Retail & E-commerce🏭 Manufacturing📡 IoT🚗 Automotive🏛️ Government & Public Sector🎓 Education📶 Telecom☁️ SaaS & Cloud🛢️ Oil & Gas⚡ Energy & Utilities🛡️ Insurance🚚 Logistics & Supply Chain✈️ Travel & Hospitality🎬 Media & Entertainment💊 Pharma & Life Sciences🏢 Real Estate & PropTech🌾 Agriculture🛫 Aviation🍔 Food Delivery & QSR🔐 Cybersecurity👥 HR & Workforce⚖️ Legal🤝 Non-profit & NGO🏅 Sports & Fitness🎮 Gaming🚀 Aerospace & Space🤖 Robotics🎨 3D Rendering & VFX🏗️ Construction🛋️ Home Interiors🌆 Urban & City Planning🧠 AI & Machine Learning📝 System Design Interview Classics📊 Data Engineering & Analytics🛠️ DevOps & SRE💹 Fintech & Capital Markets⛓️ Blockchain & Web3📣 Marketing, CRM & Support🚨 Public Safety & Emergency🚆 Railways & Public Transport🚢 Maritime, Ports & Shipping🌍 Environment, Climate & Water🛕 Religious & Spiritual Organisations🎟️ Events, Ticketing & Creators⛏️ Mining, Metals & Chemicals📋 Business Analysis & Process🎓 Student Projects (College Reports)🔬 Research & Academia

17 templates

Flowchart

GitOps Deployment Flow

How a code change reaches production with GitOps: CI builds and tests, the image tag is written to a config repo, and Argo CD syncs the cluster to match Git.

C4 ArchitecturePro

Observability Stack (Metrics, Logs, Traces)

A complete observability setup: OpenTelemetry collects metrics, logs and traces from services, which are stored separately and viewed together in Grafana with alerting.

State Machine

Incident Lifecycle

The stages of a production incident from the first alert to the postmortem, with severity escalation along the way.

Sequence

Incident Response Sequence

Who does what when production breaks: monitoring pages the on-call engineer, an incident channel is opened, customers are updated and the fix is rolled out.

DeploymentPro

Blue-Green Deployment

A blue-green setup where a new version is deployed next to the live one and traffic is switched only after it passes checks, allowing instant rollback.

Flowchart

Infrastructure as Code Pipeline

How infrastructure changes are made safely with Terraform: plan on pull request, policy checks, approval, apply and drift detection.

C4 ArchitecturePro

Internal Developer Platform

A self-service platform for developers: a Backstage portal with templates that create repos, pipelines and environments without tickets to the ops team.

Sequence

Canary Release Sequence

How a canary release sends a small share of traffic to a new version, compares its error rate with the stable version, and promotes or aborts automatically.

Network

Multi-Region Active-Active Network

A network layout for running an application in two regions at once, with global load balancing, private links between regions and replicated databases.

Sankey

Monthly Cloud Cost Breakdown

Where a company's monthly cloud bill goes, by team and by service type, to spot the biggest savings opportunities.

Mindmap

SRE Practices

The core practices of site reliability engineering, from service level objectives to toil reduction, as a map for teams starting SRE.

Flowchart

Incident Response Flow

From the first alert to the postmortem: triage by severity, assign an incident commander, mitigate, communicate and learn.

Sequence

On-call Escalation

How a page moves from the monitoring system to the primary on-call, then to the secondary and the engineering manager when nobody answers.

Flowchart

Database Failover Runbook

Step-by-step runbook for failing over from a broken primary database to a replica, with checks before and after.

Flowchart

Deployment Rollback Runbook

What to do when a release goes bad: decide quickly between a feature-flag switch-off, a rollback and a hotfix.

Flowchart

Disk Full Alert Runbook

Runbook for a disk-usage alert on a server or volume: find what is growing, free space safely, then stop it happening again.

Flowchart

TLS Certificate Expiry Runbook

Renew a TLS certificate before it expires, or recover quickly when it already has.

More DevOps & SRE templates are being added regularly.