
🧠 AI & Machine Learning · C4 Architecture · Pro
A chatbot that answers staff questions from company documents: documents are split and indexed in a vector database, and each question pulls the most relevant passages before the language model writes an answer.
C4 Architecture diagrams are part of Pro. Anyone can view this one; generating and editing it needs Pro.
Drawing diagram…
An internal RAG chatbot. Employees use a chat web app that calls a chat API. The chat API checks the user's permissions, embeds the question, searches a vector database (pgvector) for relevant passages and sends them with the question to an LLM through an LLM gateway. An ingestion service reads documents from SharePoint and Google Drive, splits them into chunks, creates embeddings and stores them in the vector database. Conversations are stored in PostgreSQL.
C4Container
title RAG Chatbot over Company Documents
Person(emp, "Employee", "Asks questions")
System_Boundary(rag, "Knowledge Assistant") {
Container(web, "Chat Web App", "React", "Chat UI with source links")
Container(api, "Chat API", "Python FastAPI", "Permissions, retrieval, prompts")
Container(ingest, "Ingestion Service", "Python", "Split, embed, index documents")
Container(gw, "LLM Gateway", "LiteLLM", "Routing, cost limits, safety filters")
ContainerDb(vec, "Vector Store", "PostgreSQL pgvector", "Chunks and embeddings")
ContainerDb(db, "Chat DB", "PostgreSQL", "Conversations and feedback")
}
System_Ext(docs, "Document Sources", "SharePoint, Google Drive")
System_Ext(llm, "LLM Provider", "Hosted language model")
Rel(emp, web, "Uses")
Rel(web, api, "Sends question", "HTTPS")
Rel(api, vec, "Similarity search")
Rel(api, gw, "Prompt with passages")
Rel(gw, llm, "Completion request")
Rel(api, db, "Stores chat")
Rel(ingest, docs, "Reads documents")
Rel(ingest, gw, "Creates embeddings")
Rel(ingest, vec, "Writes chunks")What happens when a user asks the document chatbot a question: permission check, embedding, vector search, prompt building and a cited answer.
How documents become searchable passages for an AI assistant: extraction, cleaning, chunking, embedding and indexing, with changed files re-processed automatically.
How an AI agent completes a task by planning, calling tools, checking results and asking a person to approve risky actions before it finishes.
Where each part of an MLOps setup runs: feature store, training jobs on GPU nodes, experiment tracking, model registry and automated deployment to serving.
The states a machine learning model moves through, from experiment to production to retirement, including rollback when live performance drops.
One gateway that every app in a company uses to reach language models: it applies budgets, rate limits, content filters and caching, and records usage per team.