
🛠️ DevOps & SRE · Deployment · Pro
A blue-green setup where a new version is deployed next to the live one and traffic is switched only after it passes checks, allowing instant rollback.
Deployment diagrams are part of Pro. Anyone can view this one; generating and editing it needs Pro.
Drawing diagram…
Blue-green deployment on AWS: Route 53 points to an Application Load Balancer with two target groups. Blue runs version 1.4 on an auto scaling group and serves all traffic. Green runs version 1.5 and receives only smoke-test traffic. After checks pass the listener switches to green; blue stays running for an hour for rollback. Both use the same RDS database with backward-compatible schema changes.
flowchart TB
U[Users] --> R53[Route 53]
R53 --> ALB[Application Load Balancer]
subgraph AWS[AWS Region]
subgraph BLUE[Blue - v1.4 LIVE]
B1[EC2 v1.4]
B2[EC2 v1.4]
end
subgraph GREEN[Green - v1.5 candidate]
G1[EC2 v1.5]
G2[EC2 v1.5]
end
RDS[(RDS PostgreSQL)]
end
ALB -->|100% traffic| B1
ALB -->|100% traffic| B2
ALB -.smoke tests only.-> G1
ALB -.smoke tests only.-> G2
B1 --> RDS
B2 --> RDS
G1 --> RDS
G2 --> RDSHow a code change reaches production with GitOps: CI builds and tests, the image tag is written to a config repo, and Argo CD syncs the cluster to match Git.
A complete observability setup: OpenTelemetry collects metrics, logs and traces from services, which are stored separately and viewed together in Grafana with alerting.
The stages of a production incident from the first alert to the postmortem, with severity escalation along the way.
Who does what when production breaks: monitoring pages the on-call engineer, an incident channel is opened, customers are updated and the fix is rolled out.
How infrastructure changes are made safely with Terraform: plan on pull request, policy checks, approval, apply and drift detection.
A self-service platform for developers: a Backstage portal with templates that create repos, pipelines and environments without tickets to the ops team.