
🧠 AI & Machine Learning · Database ERD
Tables that keep machine learning work reproducible: datasets and their versions, experiments, runs, metrics, and registered models with their deployments.
Drawing diagram…
Database for experiment tracking: projects have experiments; an experiment has runs; each run uses one dataset version, stores parameters and metrics; a run can produce a model version in the registry; model versions have deployments to environments.
erDiagram
PROJECT ||--o{ EXPERIMENT : has
EXPERIMENT ||--o{ RUN : contains
DATASET ||--|{ DATASET_VERSION : "versioned as"
DATASET_VERSION ||--o{ RUN : "used by"
RUN ||--o{ METRIC : logs
RUN ||--o{ PARAM : uses
RUN ||--o| MODEL_VERSION : produces
MODEL_VERSION ||--o{ DEPLOYMENT : "deployed as"
PROJECT {
int id PK
string name
string owner
}
EXPERIMENT {
int id PK
int project_id FK
string goal
}
DATASET {
int id PK
string name
}
DATASET_VERSION {
int id PK
int dataset_id FK
string hash
date created_on
}
RUN {
int id PK
int experiment_id FK
int dataset_version_id FK
string status
datetime started_at
}
METRIC {
int id PK
int run_id FK
string name
decimal value
}
PARAM {
int id PK
int run_id FK
string name
string value
}
MODEL_VERSION {
int id PK
int run_id FK
string stage
}
DEPLOYMENT {
int id PK
int model_version_id FK
string environment
date deployed_on
}A chatbot that answers staff questions from company documents: documents are split and indexed in a vector database, and each question pulls the most relevant passages before the language model writes an answer.
What happens when a user asks the document chatbot a question: permission check, embedding, vector search, prompt building and a cited answer.
How documents become searchable passages for an AI assistant: extraction, cleaning, chunking, embedding and indexing, with changed files re-processed automatically.
How an AI agent completes a task by planning, calling tools, checking results and asking a person to approve risky actions before it finishes.
Where each part of an MLOps setup runs: feature store, training jobs on GPU nodes, experiment tracking, model registry and automated deployment to serving.
The states a machine learning model moves through, from experiment to production to retirement, including rollback when live performance drops.