
🛠️ DevOps & SRE · C4 Architecture · Pro
A self-service platform for developers: a Backstage portal with templates that create repos, pipelines and environments without tickets to the ops team.
C4 Architecture diagrams are part of Pro. Anyone can view this one; generating and editing it needs Pro.
Drawing diagram…
Internal developer platform: developers use a Backstage portal with a service catalog and software templates. Creating a new service scaffolds a GitHub repo with a CI pipeline in GitHub Actions and registers it in the catalog. Crossplane provisions databases and queues from Kubernetes resources. Argo CD deploys to shared Kubernetes clusters. TechDocs publishes documentation from the repo.
C4Container
title Internal Developer Platform
Person(dev, "Developer", "Creates and runs services")
Person(plat, "Platform Team", "Maintains golden paths")
System_Boundary(idp, "Developer Platform") {
Container(bs, "Portal", "Backstage", "Catalog, templates, docs")
Container(tpl, "Software Templates", "Backstage scaffolder", "Golden-path services")
Container(xp, "Infra Provisioner", "Crossplane", "Databases, queues, buckets")
Container(argo, "Deployer", "Argo CD", "Syncs apps to clusters")
}
System_Ext(gh, "GitHub", "Repos and Actions CI")
System_Ext(k8s, "Kubernetes Clusters", "Shared runtime")
System_Ext(cloud, "Cloud Provider", "Managed services")
Rel(dev, bs, "Uses")
Rel(plat, tpl, "Maintains")
Rel(bs, tpl, "Runs")
Rel(tpl, gh, "Creates repo and pipeline")
Rel(gh, argo, "Updates manifests")
Rel(argo, k8s, "Deploys")
Rel(xp, cloud, "Provisions")
Rel(k8s, xp, "Resource claims")How a code change reaches production with GitOps: CI builds and tests, the image tag is written to a config repo, and Argo CD syncs the cluster to match Git.
A complete observability setup: OpenTelemetry collects metrics, logs and traces from services, which are stored separately and viewed together in Grafana with alerting.
The stages of a production incident from the first alert to the postmortem, with severity escalation along the way.
Who does what when production breaks: monitoring pages the on-call engineer, an incident channel is opened, customers are updated and the fix is rolled out.
A blue-green setup where a new version is deployed next to the live one and traffic is switched only after it passes checks, allowing instant rollback.
How infrastructure changes are made safely with Terraform: plan on pull request, policy checks, approval, apply and drift detection.