CloudSketch AI Logo
FeaturesTemplatesPricingEnterpriseAboutContactLog in
🌙☀️
Log inStart free
Home/Templates/AI & Machine Learning

🧠 AI & Machine Learning · C4 Architecture · Pro

RAG Chatbot over Company Documents

A chatbot that answers staff questions from company documents: documents are split and indexed in a vector database, and each question pulls the most relevant passages before the language model writes an answer.

More AI & Machine Learning templates

C4 Architecture diagrams are part of Pro. Anyone can view this one; generating and editing it needs Pro.

Drawing diagram…

What this diagram shows

  • Answers cite the passages they came from
  • Document permissions are checked before retrieval
  • The language model is called through one gateway for cost and safety limits

Prompt used

An internal RAG chatbot. Employees use a chat web app that calls a chat API. The chat API checks the user's permissions, embeds the question, searches a vector database (pgvector) for relevant passages and sends them with the question to an LLM through an LLM gateway. An ingestion service reads documents from SharePoint and Google Drive, splits them into chunks, creates embeddings and stores them in the vector database. Conversations are stored in PostgreSQL.

Mermaid code
C4Container
title RAG Chatbot over Company Documents
Person(emp, "Employee", "Asks questions")
System_Boundary(rag, "Knowledge Assistant") {
  Container(web, "Chat Web App", "React", "Chat UI with source links")
  Container(api, "Chat API", "Python FastAPI", "Permissions, retrieval, prompts")
  Container(ingest, "Ingestion Service", "Python", "Split, embed, index documents")
  Container(gw, "LLM Gateway", "LiteLLM", "Routing, cost limits, safety filters")
  ContainerDb(vec, "Vector Store", "PostgreSQL pgvector", "Chunks and embeddings")
  ContainerDb(db, "Chat DB", "PostgreSQL", "Conversations and feedback")
}
System_Ext(docs, "Document Sources", "SharePoint, Google Drive")
System_Ext(llm, "LLM Provider", "Hosted language model")
Rel(emp, web, "Uses")
Rel(web, api, "Sends question", "HTTPS")
Rel(api, vec, "Similarity search")
Rel(api, gw, "Prompt with passages")
Rel(gw, llm, "Completion request")
Rel(api, db, "Stores chat")
Rel(ingest, docs, "Reads documents")
Rel(ingest, gw, "Creates embeddings")
Rel(ingest, vec, "Writes chunks")

Related templates

SequenceAI & Machine Learning

RAG Question Answering Sequence

What happens when a user asks the document chatbot a question: permission check, embedding, vector search, prompt building and a cited answer.

Data FlowAI & Machine Learning

Document Ingestion for Vector Search

How documents become searchable passages for an AI assistant: extraction, cleaning, chunking, embedding and indexing, with changed files re-processed automatically.

FlowchartAI & Machine Learning

AI Agent Tool-Use Workflow

How an AI agent completes a task by planning, calling tools, checking results and asking a person to approve risky actions before it finishes.

DeploymentProAI & Machine Learning

MLOps Model Training Pipeline

Where each part of an MLOps setup runs: feature store, training jobs on GPU nodes, experiment tracking, model registry and automated deployment to serving.

State MachineAI & Machine Learning

ML Model Lifecycle

The states a machine learning model moves through, from experiment to production to retirement, including rollback when live performance drops.

C4 ArchitectureProAI & Machine Learning

LLM Gateway with Cost and Safety Controls

One gateway that every app in a company uses to reach language models: it applies budgets, rate limits, content filters and caching, and records usage per team.