
🧠 AI & Machine Learning · Sequence
What happens when a user asks the document chatbot a question: permission check, embedding, vector search, prompt building and a cited answer.
Drawing diagram…
Sequence for a RAG chatbot: the user sends a question to the chat API. The API gets the user's allowed document groups from the identity service, asks the embedding model for a vector, searches the vector store with a permission filter, re-ranks the top 20 passages to keep the best 5, builds a prompt and calls the LLM, then streams the answer with citations back to the user and logs the exchange.
sequenceDiagram actor U as User participant API as Chat API participant ID as Identity Service participant EMB as Embedding Model participant VS as Vector Store participant RR as Re-ranker participant LLM as Language Model participant LOG as Chat Log U->>API: Ask question API->>ID: Get allowed document groups ID-->>API: Groups API->>EMB: Embed question EMB-->>API: Vector API->>VS: Top 20 passages in allowed groups VS-->>API: Passages API->>RR: Re-rank passages RR-->>API: Best 5 passages API->>LLM: Prompt with question and passages LLM-->>API: Answer tokens API-->>U: Streamed answer with sources API->>LOG: Save question, answer, sources
A chatbot that answers staff questions from company documents: documents are split and indexed in a vector database, and each question pulls the most relevant passages before the language model writes an answer.
How documents become searchable passages for an AI assistant: extraction, cleaning, chunking, embedding and indexing, with changed files re-processed automatically.
How an AI agent completes a task by planning, calling tools, checking results and asking a person to approve risky actions before it finishes.
Where each part of an MLOps setup runs: feature store, training jobs on GPU nodes, experiment tracking, model registry and automated deployment to serving.
The states a machine learning model moves through, from experiment to production to retirement, including rollback when live performance drops.
One gateway that every app in a company uses to reach language models: it applies budgets, rate limits, content filters and caching, and records usage per team.