
🧠 AI & Machine Learning · Deployment · Pro
Where each part of an MLOps setup runs: feature store, training jobs on GPU nodes, experiment tracking, model registry and automated deployment to serving.
Deployment diagrams are part of Pro. Anyone can view this one; generating and editing it needs Pro.
Drawing diagram…
Deployment of an MLOps platform on Google Cloud: data lands in BigQuery and a Feast feature store. Kubeflow Pipelines on GKE runs training jobs on a GPU node pool, logs runs to MLflow, and registers models in the MLflow model registry. After approval, Argo CD deploys the model to a KServe serving cluster behind a load balancer. Models and artifacts are stored in Cloud Storage.
flowchart TB
DS[Data Scientists] --> KF
subgraph GCP[Google Cloud]
BQ[(BigQuery)]
FS[(Feast Feature Store)]
GCS[(Cloud Storage - artifacts)]
subgraph TRAIN[GKE Training Cluster]
KF[Kubeflow Pipelines]
subgraph GPU[GPU Node Pool]
JOB[Training Jobs]
end
MLF[MLflow Tracking and Registry]
end
subgraph SERVE[GKE Serving Cluster]
ARGO[Argo CD]
KS[KServe Model Pods]
end
LB[Load Balancer]
end
BQ --> FS
KF --> JOB
FS --> JOB
JOB --> MLF
JOB --> GCS
MLF -->|Approved version| ARGO
ARGO --> KS
GCS --> KS
LB --> KS
APPS[Client Apps] --> LBA chatbot that answers staff questions from company documents: documents are split and indexed in a vector database, and each question pulls the most relevant passages before the language model writes an answer.
What happens when a user asks the document chatbot a question: permission check, embedding, vector search, prompt building and a cited answer.
How documents become searchable passages for an AI assistant: extraction, cleaning, chunking, embedding and indexing, with changed files re-processed automatically.
How an AI agent completes a task by planning, calling tools, checking results and asking a person to approve risky actions before it finishes.
The states a machine learning model moves through, from experiment to production to retirement, including rollback when live performance drops.
One gateway that every app in a company uses to reach language models: it applies budgets, rate limits, content filters and caching, and records usage per team.