
🛠️ DevOps & SRE
GitOps, observability, incident response, on-call runbooks, release strategies and infrastructure as code. Open any template to see the finished diagram, then use it as the starting point for your own in the Studio.
17 templates
How a code change reaches production with GitOps: CI builds and tests, the image tag is written to a config repo, and Argo CD syncs the cluster to match Git.
A complete observability setup: OpenTelemetry collects metrics, logs and traces from services, which are stored separately and viewed together in Grafana with alerting.
The stages of a production incident from the first alert to the postmortem, with severity escalation along the way.
Who does what when production breaks: monitoring pages the on-call engineer, an incident channel is opened, customers are updated and the fix is rolled out.
A blue-green setup where a new version is deployed next to the live one and traffic is switched only after it passes checks, allowing instant rollback.
How infrastructure changes are made safely with Terraform: plan on pull request, policy checks, approval, apply and drift detection.
A self-service platform for developers: a Backstage portal with templates that create repos, pipelines and environments without tickets to the ops team.
How a canary release sends a small share of traffic to a new version, compares its error rate with the stable version, and promotes or aborts automatically.
A network layout for running an application in two regions at once, with global load balancing, private links between regions and replicated databases.
Where a company's monthly cloud bill goes, by team and by service type, to spot the biggest savings opportunities.
The core practices of site reliability engineering, from service level objectives to toil reduction, as a map for teams starting SRE.
From the first alert to the postmortem: triage by severity, assign an incident commander, mitigate, communicate and learn.
How a page moves from the monitoring system to the primary on-call, then to the secondary and the engineering manager when nobody answers.
Step-by-step runbook for failing over from a broken primary database to a replica, with checks before and after.
What to do when a release goes bad: decide quickly between a feature-flag switch-off, a rollback and a hotfix.
Runbook for a disk-usage alert on a server or volume: find what is growing, free space safely, then stop it happening again.
Renew a TLS certificate before it expires, or recover quickly when it already has.
More DevOps & SRE templates are being added regularly.