
🛠️ DevOps & SRE · Mindmap
The core practices of site reliability engineering, from service level objectives to toil reduction, as a map for teams starting SRE.
Drawing diagram…
Mind map of SRE practices: service levels (SLIs, SLOs, error budgets), monitoring (golden signals, alerting on symptoms, dashboards), incident management (on-call rotation, runbooks, postmortems), release engineering (canary, feature flags, rollback), capacity planning (load tests, forecasting) and toil reduction (automation, self-service).
mindmap
root((SRE Practices))
Service levels
SLIs
SLOs
Error budgets
Monitoring
Golden signals
Symptom-based alerts
Dashboards
Incidents
On-call rotation
Runbooks
Blameless postmortems
Releases
Canary releases
Feature flags
Fast rollback
Capacity
Load testing
Forecasting
Toil
Automation
Self-service toolsHow a code change reaches production with GitOps: CI builds and tests, the image tag is written to a config repo, and Argo CD syncs the cluster to match Git.
A complete observability setup: OpenTelemetry collects metrics, logs and traces from services, which are stored separately and viewed together in Grafana with alerting.
The stages of a production incident from the first alert to the postmortem, with severity escalation along the way.
Who does what when production breaks: monitoring pages the on-call engineer, an incident channel is opened, customers are updated and the fix is rolled out.
A blue-green setup where a new version is deployed next to the live one and traffic is switched only after it passes checks, allowing instant rollback.
How infrastructure changes are made safely with Terraform: plan on pull request, policy checks, approval, apply and drift detection.