
🛠️ DevOps & SRE · Network
A network layout for running an application in two regions at once, with global load balancing, private links between regions and replicated databases.
Drawing diagram…
Network for a multi-region active-active setup: a global load balancer with health checks routes users to Mumbai or Singapore. Each region has a VPC with public subnets for load balancers and private subnets for app servers and databases. The two VPCs are connected with a transit gateway peering. Aurora Global Database replicates between regions and Redis uses active-active replication.
flowchart LR
USERS((Users)) --- GLB[Global Load Balancer]
subgraph MUM[Region Mumbai - VPC 10.0.0.0/16]
ALB1[Public ALB]
APP1[App Servers - private subnet]
DB1[(Aurora Primary)]
RC1[(Redis)]
end
subgraph SIN[Region Singapore - VPC 10.1.0.0/16]
ALB2[Public ALB]
APP2[App Servers - private subnet]
DB2[(Aurora Replica)]
RC2[(Redis)]
end
GLB --- ALB1
GLB --- ALB2
ALB1 --- APP1
ALB2 --- APP2
APP1 --- DB1
APP1 --- RC1
APP2 --- DB2
APP2 --- RC2
TGW1[Transit Gateway] --- TGW2[Transit Gateway]
APP1 --- TGW1
APP2 --- TGW2
DB1 -.replication.- DB2
RC1 -.active-active.- RC2How a code change reaches production with GitOps: CI builds and tests, the image tag is written to a config repo, and Argo CD syncs the cluster to match Git.
A complete observability setup: OpenTelemetry collects metrics, logs and traces from services, which are stored separately and viewed together in Grafana with alerting.
The stages of a production incident from the first alert to the postmortem, with severity escalation along the way.
Who does what when production breaks: monitoring pages the on-call engineer, an incident channel is opened, customers are updated and the fix is rolled out.
A blue-green setup where a new version is deployed next to the live one and traffic is switched only after it passes checks, allowing instant rollback.
How infrastructure changes are made safely with Terraform: plan on pull request, policy checks, approval, apply and drift detection.