CloudSketch AI Logo
FeaturesTemplatesPricingEnterpriseAboutContactLog in
🌙☀️
Log inStart free
Home/Templates/AI & Machine Learning

🧠 AI & Machine Learning · Deployment · Pro

GPU Inference Cluster Deployment

How a self-hosted language model is deployed for many users: autoscaled GPU pods behind a gateway, with model weights cached close to the GPUs.

More AI & Machine Learning templates

Deployment diagrams are part of Pro. Anyone can view this one; generating and editing it needs Pro.

Drawing diagram…

What this diagram shows

  • GPU pods scale with the request queue length
  • Model weights are cached on node-local disks
  • Requests are batched to use the GPUs fully

Prompt used

Deployment for self-hosted LLM inference on AWS: users call an API gateway, which forwards to an inference router on EKS. vLLM pods run on a GPU node group (g5 instances) and scale with KEDA based on queue length. Model weights come from S3 and are cached on NVMe. Prometheus and Grafana watch GPU usage and latency.

Mermaid code
flowchart TB
  U[Client Apps] --> APIGW[API Gateway]
  subgraph AWS[AWS Region us-east-1]
    subgraph EKS[EKS Cluster]
      R[Inference Router]
      subgraph GPUNG[GPU Node Group - g5]
        V1[vLLM Pod 1]
        V2[vLLM Pod 2]
        V3[vLLM Pod 3]
        NV[(Local NVMe weight cache)]
      end
      KEDA[KEDA Autoscaler]
      PROM[Prometheus]
      GRAF[Grafana]
    end
    S3[(S3 Model Weights)]
  end
  APIGW --> R
  R --> V1
  R --> V2
  R --> V3
  S3 --> NV
  NV --> V1
  NV --> V2
  NV --> V3
  KEDA -.scales.-> GPUNG
  PROM -.scrapes.-> R
  GRAF --> PROM

Related templates

C4 ArchitectureProAI & Machine Learning

RAG Chatbot over Company Documents

A chatbot that answers staff questions from company documents: documents are split and indexed in a vector database, and each question pulls the most relevant passages before the language model writes an answer.

SequenceAI & Machine Learning

RAG Question Answering Sequence

What happens when a user asks the document chatbot a question: permission check, embedding, vector search, prompt building and a cited answer.

Data FlowAI & Machine Learning

Document Ingestion for Vector Search

How documents become searchable passages for an AI assistant: extraction, cleaning, chunking, embedding and indexing, with changed files re-processed automatically.

FlowchartAI & Machine Learning

AI Agent Tool-Use Workflow

How an AI agent completes a task by planning, calling tools, checking results and asking a person to approve risky actions before it finishes.

DeploymentProAI & Machine Learning

MLOps Model Training Pipeline

Where each part of an MLOps setup runs: feature store, training jobs on GPU nodes, experiment tracking, model registry and automated deployment to serving.

State MachineAI & Machine Learning

ML Model Lifecycle

The states a machine learning model moves through, from experiment to production to retirement, including rollback when live performance drops.