
🧠 AI & Machine Learning · Data Flow
How documents become searchable passages for an AI assistant: extraction, cleaning, chunking, embedding and indexing, with changed files re-processed automatically.
Drawing diagram…
Data flow for document ingestion: a connector detects new or changed files in SharePoint, Google Drive and Confluence and sends them to a text extractor (OCR for scanned PDFs). Clean text is split into chunks of about 500 tokens with overlap, sent to an embedding model, and stored with metadata (source, page, access group) in a vector database. A metadata store tracks file versions so unchanged files are skipped.
flowchart LR SRC[(SharePoint, Drive, Confluence)] -->|New or changed files| CON[Connector] META[(File Version Store)] -->|Last seen versions| CON CON -->|Files| EXT[Text Extractor and OCR] EXT -->|Clean text| CH[Chunker - 500 tokens] CH -->|Chunks| EMB[Embedding Model] EMB -->|Vectors| IDX[Indexer] CH -->|Source, page, access group| IDX IDX -->|Chunks and vectors| VDB[(Vector Database)] CON -->|Processed versions| META
A chatbot that answers staff questions from company documents: documents are split and indexed in a vector database, and each question pulls the most relevant passages before the language model writes an answer.
What happens when a user asks the document chatbot a question: permission check, embedding, vector search, prompt building and a cited answer.
How an AI agent completes a task by planning, calling tools, checking results and asking a person to approve risky actions before it finishes.
Where each part of an MLOps setup runs: feature store, training jobs on GPU nodes, experiment tracking, model registry and automated deployment to serving.
The states a machine learning model moves through, from experiment to production to retirement, including rollback when live performance drops.
One gateway that every app in a company uses to reach language models: it applies budgets, rate limits, content filters and caching, and records usage per team.