
🛠️ DevOps & SRE · Sequence
How a page moves from the monitoring system to the primary on-call, then to the secondary and the engineering manager when nobody answers.
Drawing diagram…
Sequence diagram of on-call escalation: Monitoring sends an alert to the Paging service. The paging service pages the Primary on-call. If there is no acknowledgement in 5 minutes it pages the Secondary on-call, and after 10 more minutes the Engineering manager. The engineer who acknowledges opens an incident in the Incident channel, the Incident commander joins, posts to the Status page and informs Stakeholders. When resolved the commander closes the incident and the paging service stops.
sequenceDiagram
participant Mon as Monitoring
participant Pg as Paging service
participant P as Primary on-call
participant S as Secondary on-call
participant EM as Engineering manager
participant IC as Incident commander
participant SP as Status page
Mon->>Pg: Alert (error rate above threshold)
Pg->>P: Page
alt No acknowledgement in 5 min
Pg->>S: Page secondary
alt No acknowledgement in 10 min
Pg->>EM: Page manager
end
end
P->>Pg: Acknowledge (stops escalation)
P->>IC: Open incident, request commander
IC->>SP: Post "Investigating"
loop Every 30 minutes
IC->>SP: Status update
end
P->>IC: Mitigated
IC->>SP: Post "Resolved"
IC->>Pg: Close incidentHow a code change reaches production with GitOps: CI builds and tests, the image tag is written to a config repo, and Argo CD syncs the cluster to match Git.
A complete observability setup: OpenTelemetry collects metrics, logs and traces from services, which are stored separately and viewed together in Grafana with alerting.
The stages of a production incident from the first alert to the postmortem, with severity escalation along the way.
Who does what when production breaks: monitoring pages the on-call engineer, an incident channel is opened, customers are updated and the fix is rolled out.
A blue-green setup where a new version is deployed next to the live one and traffic is switched only after it passes checks, allowing instant rollback.
How infrastructure changes are made safely with Terraform: plan on pull request, policy checks, approval, apply and drift detection.