CloudSketch AI Logo
FeaturesTemplatesPricingEnterpriseAboutContactLog in
🌙☀️
Log inStart free
Home/Templates/AI & Machine Learning

🧠 AI & Machine Learning · Data Flow

Document Ingestion for Vector Search

How documents become searchable passages for an AI assistant: extraction, cleaning, chunking, embedding and indexing, with changed files re-processed automatically.

More AI & Machine Learning templates

Drawing diagram…

What this diagram shows

  • Only new or changed files are re-processed
  • Scanned PDFs go through OCR first
  • Each chunk keeps its source and access group

Prompt used

Data flow for document ingestion: a connector detects new or changed files in SharePoint, Google Drive and Confluence and sends them to a text extractor (OCR for scanned PDFs). Clean text is split into chunks of about 500 tokens with overlap, sent to an embedding model, and stored with metadata (source, page, access group) in a vector database. A metadata store tracks file versions so unchanged files are skipped.

Mermaid code
flowchart LR
  SRC[(SharePoint, Drive, Confluence)] -->|New or changed files| CON[Connector]
  META[(File Version Store)] -->|Last seen versions| CON
  CON -->|Files| EXT[Text Extractor and OCR]
  EXT -->|Clean text| CH[Chunker - 500 tokens]
  CH -->|Chunks| EMB[Embedding Model]
  EMB -->|Vectors| IDX[Indexer]
  CH -->|Source, page, access group| IDX
  IDX -->|Chunks and vectors| VDB[(Vector Database)]
  CON -->|Processed versions| META

Related templates

C4 ArchitectureProAI & Machine Learning

RAG Chatbot over Company Documents

A chatbot that answers staff questions from company documents: documents are split and indexed in a vector database, and each question pulls the most relevant passages before the language model writes an answer.

SequenceAI & Machine Learning

RAG Question Answering Sequence

What happens when a user asks the document chatbot a question: permission check, embedding, vector search, prompt building and a cited answer.

FlowchartAI & Machine Learning

AI Agent Tool-Use Workflow

How an AI agent completes a task by planning, calling tools, checking results and asking a person to approve risky actions before it finishes.

DeploymentProAI & Machine Learning

MLOps Model Training Pipeline

Where each part of an MLOps setup runs: feature store, training jobs on GPU nodes, experiment tracking, model registry and automated deployment to serving.

State MachineAI & Machine Learning

ML Model Lifecycle

The states a machine learning model moves through, from experiment to production to retirement, including rollback when live performance drops.

C4 ArchitectureProAI & Machine Learning

LLM Gateway with Cost and Safety Controls

One gateway that every app in a company uses to reach language models: it applies budgets, rate limits, content filters and caching, and records usage per team.