
🛠️ DevOps & SRE · Sequence
How a canary release sends a small share of traffic to a new version, compares its error rate with the stable version, and promotes or aborts automatically.
Drawing diagram…
Canary release with Argo Rollouts: the controller deploys version 2 pods and tells the service mesh to send 5 percent of traffic to them. It queries Prometheus to compare error rate and latency with version 1. If metrics are healthy it increases to 25, 50 and 100 percent. If error rate rises above 1 percent it sets canary traffic to zero and scales version 2 down.
sequenceDiagram
participant AR as Argo Rollouts
participant K as Kubernetes
participant M as Service Mesh
participant P as Prometheus
AR->>K: Create v2 pods
AR->>M: Send 5% to v2
loop Each step: 5, 25, 50%
AR->>P: Compare v2 and v1 errors and latency
alt Healthy
P-->>AR: v2 within limits
AR->>M: Increase v2 share
else Error rate above 1%
P-->>AR: v2 failing
AR->>M: Send 0% to v2
AR->>K: Scale v2 down
end
end
AR->>M: Send 100% to v2
AR->>K: Remove v1 podsHow a code change reaches production with GitOps: CI builds and tests, the image tag is written to a config repo, and Argo CD syncs the cluster to match Git.
A complete observability setup: OpenTelemetry collects metrics, logs and traces from services, which are stored separately and viewed together in Grafana with alerting.
The stages of a production incident from the first alert to the postmortem, with severity escalation along the way.
Who does what when production breaks: monitoring pages the on-call engineer, an incident channel is opened, customers are updated and the fix is rolled out.
A blue-green setup where a new version is deployed next to the live one and traffic is switched only after it passes checks, allowing instant rollback.
How infrastructure changes are made safely with Terraform: plan on pull request, policy checks, approval, apply and drift detection.