CloudSketch AI Logo
FeaturesTemplatesPricingEnterpriseAboutContactLog in
🌙☀️
Log inStart free
Home/Templates/AI & Machine Learning

🧠 AI & Machine Learning · C4 Architecture · Pro

LLM Gateway with Cost and Safety Controls

One gateway that every app in a company uses to reach language models: it applies budgets, rate limits, content filters and caching, and records usage per team.

More AI & Machine Learning templates

C4 Architecture diagrams are part of Pro. Anyone can view this one; generating and editing it needs Pro.

Drawing diagram…

What this diagram shows

  • Teams get monthly token budgets
  • Prompts and answers pass through safety filters
  • Repeated prompts are served from a cache

Prompt used

An enterprise LLM gateway: internal apps call the gateway API with a team key. The gateway checks the team's budget and rate limit in Redis, removes personal data with a PII filter, checks a semantic cache, and routes the request to OpenAI, Anthropic or a self-hosted model depending on policy. Usage and cost are written to a usage database and shown on an admin dashboard.

Mermaid code
C4Container
title LLM Gateway
Person(dev, "App Team", "Builds AI features")
Person(admin, "Platform Admin", "Sets budgets and policies")
System_Boundary(gw, "LLM Gateway") {
  Container(api, "Gateway API", "Go", "Auth, routing, streaming")
  Container(pii, "PII Filter", "Python", "Masks personal data")
  Container(dash, "Admin Dashboard", "React", "Budgets, usage, policies")
  ContainerDb(redis, "Limits and Cache", "Redis", "Budgets, rate limits, semantic cache")
  ContainerDb(usage, "Usage DB", "ClickHouse", "Tokens and cost per team")
}
System_Ext(openai, "Hosted Model A", "Commercial LLM API")
System_Ext(anth, "Hosted Model B", "Commercial LLM API")
System_Ext(self, "Self-hosted Model", "vLLM on GPUs")
Rel(dev, api, "Calls with team key", "HTTPS")
Rel(api, redis, "Check budget and cache")
Rel(api, pii, "Filter prompt")
Rel(api, openai, "Routes")
Rel(api, anth, "Routes")
Rel(api, self, "Routes")
Rel(api, usage, "Writes usage")
Rel(admin, dash, "Uses")
Rel(dash, usage, "Reads")

Related templates

C4 ArchitectureProAI & Machine Learning

RAG Chatbot over Company Documents

A chatbot that answers staff questions from company documents: documents are split and indexed in a vector database, and each question pulls the most relevant passages before the language model writes an answer.

SequenceAI & Machine Learning

RAG Question Answering Sequence

What happens when a user asks the document chatbot a question: permission check, embedding, vector search, prompt building and a cited answer.

Data FlowAI & Machine Learning

Document Ingestion for Vector Search

How documents become searchable passages for an AI assistant: extraction, cleaning, chunking, embedding and indexing, with changed files re-processed automatically.

FlowchartAI & Machine Learning

AI Agent Tool-Use Workflow

How an AI agent completes a task by planning, calling tools, checking results and asking a person to approve risky actions before it finishes.

DeploymentProAI & Machine Learning

MLOps Model Training Pipeline

Where each part of an MLOps setup runs: feature store, training jobs on GPU nodes, experiment tracking, model registry and automated deployment to serving.

State MachineAI & Machine Learning

ML Model Lifecycle

The states a machine learning model moves through, from experiment to production to retirement, including rollback when live performance drops.