Prompt
How do I set up an experiment tracker for comparing training runs and managing model artifacts?
Latest observation
A good experiment tracker usually gives you three things:
- Run comparison — metrics, plots, configs, and tags across training runs
- Artifact management — model checkpoints, datasets, feature files, plots, and reports
- Reproducibility — code version, environment, and parameters tied to each run
Here’s a practical way to set one up.
1) Pick a tracker
Common options:
- Weights & Biases (W&B) — easy, feature-rich, great UI
- MLflow — open-source, flexible, good for self-hosting
- Comet — polished UI, good collaboration features
- Neptune — strong for experiment metadata
- TensorBoard — simple metric visualization, less complete for artifact management
If you want a solid general-purpose setup, MLflow is a good default. If you want the best UX and collaboration, W&B is often the quickest win.
2) Decide what you want to log
At minimum, log:
Run metadata
- model name
- dataset/version
- training script or git commit hash
- date/time
- user/machine info
- random seed
Hyperparameters
- learning rate
- batch size
- optimizer
- architecture settings
- augmentation settings
Metrics
- training loss
- validation loss
- accuracy / F1 / AUC
- per-class metrics
- learning curves
Artifacts
- best model checkpoint
- final model
- tokenizer / preprocessing objects
- confusion matrix
- sample predictions
- config files
- training logs
Environment
- Python version
- package versions
- CUDA/cuDNN versions
- GPU type
3) Set up MLflow example
Install
pip install mlflow
Start a local tracking server
For a quick local setup:
mlflow ui --backend-store-uri ./mlruns
Then open:
http://127.0.0.1:5000
For more durable storage, you can point it to:
- a database for metadata
- S3/GCS/Azure blob for artifacts
4) Instrument your training code
Example in Python:
import mlflow
import mlflow.pytorch
mlflow.set_experiment("image-classification")
with mlflow.start_run():
mlflow.log_param("learning_rate", 1e-3)
mlflow.log_param("batch_size", 64)
mlflow.log_param("epochs", 20)
for epoch in range(20):
train_loss = 0.42
val_loss = 0.38
val_acc = 0.87
mlflow.log_metric("train_loss", train_loss, step=epoch)
mlflow.log_metric("val_loss", val_loss, step=epoch)
mlflow.log_metric("val_acc", val_acc, step=epoch)
# Save model artifact
mlflow.pytorch.log_model(model, artifact_path="model")
# Save extra files
mlflow.log_artifact("confusion_matrix.png")
mlflow.log_artifact("config.yaml")
5) Track artifacts properly
Good artifact organization helps a lot later.
Suggested structure:
model/— serialized modelmetrics/— charts, summariespredictions/— CSVs or JSONs with outputsconfigs/— run configsdata_samples/— sample inputs/outputs used for debugging
If your tracker supports it, log:
- the best checkpoint
- the final checkpoint
- a model signature or input schema
- a requirements file or environment snapshot
6) Make comparison easy
To compare training runs effectively:
- use consistent metric names
- log the same metrics at the same frequency
- group runs by:
- model family
- dataset version
- experiment name
- tag runs with useful labels:
baselineablation_lrseed_42augmented
Then compare:
- best validation score
- convergence speed
- overfitting gap
- inference latency
- model size
7) Version code, data, and configs
To make results reproducible, pair experiment tracking with versioning:
- Git for code
- DVC or dataset versioning for data
- YAML/JSON config files for parameters
- environment lock files:
requirements.txtpoetry.lockconda.yaml
Log the commit hash in every run:
import subprocess
commit = subprocess.check_output(["git", "rev-parse", "HEAD"]).decode().strip()
mlflow.log_param("git_commit", commit)
8) Add model registry / promotion workflow
If you want to manage artifacts beyond just logging them, use a registry:
- Staging
- Production
- Archived
Workflow:
- train several runs
- compare metrics
- register the best artifact
- promote it after validation
- keep older versions for rollback
MLflow Model Registry and W&B Artifacts both support this well.
9) Best practices
- Log early and often
- Keep run names human-readable
- Use tags for filtering
- Save the exact config used for each run
- Don’t store huge raw datasets in the tracker if object storage is better
- Track seeds to make comparisons fair
- Log evaluation on a fixed validation/test split
- Keep artifact names deterministic
10) Minimal recommended setup
If you want something practical and not overcomplicated:
- MLflow for experiment tracking + model registry
- Git for code
- DVC or object storage for datasets
- S3/GCS/local storage for large artifacts
- YAML configs for run parameters
If you want, I can give you:
- a complete MLflow starter template for PyTorch or TensorFlow, or
- a Weights & Biases setup with artifact tracking and run comparison.