
📝 System Design Interview Classics · Deployment · Pro
A notification system that sends email, SMS and push messages at scale: one API, a queue per channel, user preferences, and retries.
Deployment diagrams are part of Pro. Anyone can view this one; generating and editing it needs Pro.
Drawing diagram…
Notification service deployment: internal services call the notification API on Kubernetes. The API checks user preferences in PostgreSQL, renders templates and puts messages on per-channel Kafka topics. Email, SMS and push workers consume their topics and call SendGrid, an SMS gateway and Firebase Cloud Messaging. Failures go to a retry topic and then a dead-letter topic.
flowchart TB
SVC[Internal Services] --> API
subgraph K8S[Kubernetes Cluster]
API[Notification API]
TPL[Template Renderer]
subgraph WORKERS[Workers]
EW[Email Worker]
SW[SMS Worker]
PW[Push Worker]
end
end
subgraph KAFKA[Kafka]
TE[email topic]
TS[sms topic]
TP[push topic]
TR[retry topic]
DLQ[dead-letter topic]
end
PREF[(Preferences DB - PostgreSQL)]
API --> PREF
API --> TPL
API --> TE
API --> TS
API --> TP
TE --> EW
TS --> SW
TP --> PW
EW --> SG[SendGrid]
SW --> SMS[SMS Gateway]
PW --> FCM[Firebase Cloud Messaging]
EW -.failure.-> TR
SW -.failure.-> TR
PW -.failure.-> TR
TR -.after 5 tries.-> DLQThe classic interview design for a service like bit.ly: short code generation, fast redirects from a cache, and click analytics processed separately.
How a one-to-one chat message is delivered in a WhatsApp-style system, with sent, delivered and read ticks, and push notifications when the receiver is offline.
How posts reach followers' feeds: fan-out on write for normal users, fan-out on read for celebrities, and a ranked feed built from a cache.
The core of a ride-hailing app: drivers stream their locations, a geo index finds nearby drivers, and a matching service offers the trip to the best one.
How a token bucket rate limiter decides whether to let an API request through, using Redis so all servers share the same counts.
How search suggestions appear as you type: a prefix index built offline from past searches, served from memory, and refreshed regularly.