
🛠️ DevOps & SRE · Sequence
Who does what when production breaks: monitoring pages the on-call engineer, an incident channel is opened, customers are updated and the fix is rolled out.
Drawing diagram…
Incident response sequence: monitoring detects a spike in errors and PagerDuty pages the on-call engineer, who acknowledges and opens a Slack incident channel. The incident commander updates the status page. The engineer checks dashboards, finds the latest deploy caused it and rolls back through the CD tool. Error rates recover, the status page is updated to resolved and a postmortem is scheduled.
sequenceDiagram participant MON as Monitoring participant PD as PagerDuty actor OC as On-call Engineer participant SL as Slack Incident Channel actor IC as Incident Commander participant SP as Status Page participant CD as CD Tool MON->>PD: Error rate above 5% PD->>OC: Page OC->>PD: Acknowledge OC->>SL: Open incident channel SL->>IC: Commander joins IC->>SP: Investigating - checkout errors OC->>MON: Check dashboards MON-->>OC: Started after 14:05 deploy OC->>CD: Roll back to previous version CD-->>OC: Rollback done MON-->>SL: Error rate normal IC->>SP: Resolved IC->>SL: Schedule postmortem
How a code change reaches production with GitOps: CI builds and tests, the image tag is written to a config repo, and Argo CD syncs the cluster to match Git.
A complete observability setup: OpenTelemetry collects metrics, logs and traces from services, which are stored separately and viewed together in Grafana with alerting.
The stages of a production incident from the first alert to the postmortem, with severity escalation along the way.
A blue-green setup where a new version is deployed next to the live one and traffic is switched only after it passes checks, allowing instant rollback.
How infrastructure changes are made safely with Terraform: plan on pull request, policy checks, approval, apply and drift detection.
A self-service platform for developers: a Backstage portal with templates that create repos, pipelines and environments without tickets to the ops team.