
🧠 AI & Machine Learning · Deployment · Pro
How a self-hosted language model is deployed for many users: autoscaled GPU pods behind a gateway, with model weights cached close to the GPUs.
Deployment diagrams are part of Pro. Anyone can view this one; generating and editing it needs Pro.
Drawing diagram…
Deployment for self-hosted LLM inference on AWS: users call an API gateway, which forwards to an inference router on EKS. vLLM pods run on a GPU node group (g5 instances) and scale with KEDA based on queue length. Model weights come from S3 and are cached on NVMe. Prometheus and Grafana watch GPU usage and latency.
flowchart TB
U[Client Apps] --> APIGW[API Gateway]
subgraph AWS[AWS Region us-east-1]
subgraph EKS[EKS Cluster]
R[Inference Router]
subgraph GPUNG[GPU Node Group - g5]
V1[vLLM Pod 1]
V2[vLLM Pod 2]
V3[vLLM Pod 3]
NV[(Local NVMe weight cache)]
end
KEDA[KEDA Autoscaler]
PROM[Prometheus]
GRAF[Grafana]
end
S3[(S3 Model Weights)]
end
APIGW --> R
R --> V1
R --> V2
R --> V3
S3 --> NV
NV --> V1
NV --> V2
NV --> V3
KEDA -.scales.-> GPUNG
PROM -.scrapes.-> R
GRAF --> PROMA chatbot that answers staff questions from company documents: documents are split and indexed in a vector database, and each question pulls the most relevant passages before the language model writes an answer.
What happens when a user asks the document chatbot a question: permission check, embedding, vector search, prompt building and a cited answer.
How documents become searchable passages for an AI assistant: extraction, cleaning, chunking, embedding and indexing, with changed files re-processed automatically.
How an AI agent completes a task by planning, calling tools, checking results and asking a person to approve risky actions before it finishes.
Where each part of an MLOps setup runs: feature store, training jobs on GPU nodes, experiment tracking, model registry and automated deployment to serving.
The states a machine learning model moves through, from experiment to production to retirement, including rollback when live performance drops.