Prompt
How do I set up model versioning and rollback for production serving?
Latest observation
A good production model versioning + rollback setup usually has 4 pieces:
- Immutable model artifacts
- A registry with metadata and stages
- Deployment that points to a specific version
- A safe rollback mechanism with traffic control
Here’s a practical blueprint.
1) Version every model artifact immutably
Treat each trained model as a build artifact, not as “the model name.”
Store:
- model binary / weights
- preprocessing artifacts
- feature schema
- training code version
- training data snapshot or data version
- hyperparameters
- evaluation metrics
- environment info
Use a unique version identifier such as:
- semantic version:
reco-model-v1.4.2 - timestamped build ID:
2026-01-15_1830 - Git SHA + pipeline run ID:
model-8f3a2c1
Best practice:
- never overwrite a deployed artifact
- every retrain creates a new version
- artifacts should be content-addressed or stored in immutable object storage
2) Use a model registry
A registry gives you a source of truth for:
- model versions
- metadata
- lifecycle stage
- approval status
Common stages:
- Development
- Staging
- Production
- Archived/Deprecated
A version record should include:
- model ID
- version number
- training run ID
- metrics
- owner
- approval state
- linked artifact location
- serving config
- compatibility notes
Example:
| Model | Version | Stage | AUC | Data Version | Artifact |
|---|---|---|---|---|---|
| fraud-detector | 12 | Production | 0.981 | 2026-01-10 | s3://.../model-12/ |
| fraud-detector | 13 | Staging | 0.984 | 2026-01-14 | s3://.../model-13/ |
Tools:
- MLflow Model Registry
- SageMaker Model Registry
- Vertex AI Model Registry
- custom registry in DB + object storage
- BentoML + external registry
- Kubeflow metadata stores
3) Promote versions through environments
Use a release flow like:
Train → Validate → Staging → Canary → Production
In each step:
- run offline evaluation on fixed benchmark sets
- validate schema and feature compatibility
- check calibration, latency, memory use
- compare against current prod model
- run integration tests
- verify business metrics if possible
Typical promotion rules:
- only models above threshold metrics can move to staging
- only models approved by CI/CD or a human gate can move to prod
- require canary success before full rollout
4) Deploy by reference, not by copying code
Your inference service should load a model based on a config or version reference, e.g.:
model_name = fraud-detectormodel_version = 12
This lets you switch versions without rebuilding the app.
Deployment methods:
- Blue/green: run old and new versions side by side, switch all traffic at once
- Canary: route 1–10% traffic to new version, then gradually increase
- Shadow: send a copy of traffic to new version but don’t serve its response
- A/B testing: split users into cohorts and compare metrics
For rollback safety, blue/green and canary are the most common.
5) Make rollback a first-class operation
Rollback should be:
- fast
- automated
- tested
- version-specific
Rollback patterns
A. Instant version pointer rollback
If your service resolves a version alias like production -> model-13, rollback means repointing it:
production -> model-12
This is the simplest and often best option.
B. Redeploy last known good version
If the deployment is immutable, redeploy the previous artifact:
- restore container config
- load model v12
- restart serving pods
C. Traffic rollback
If using service mesh / load balancer routing:
- reduce new model traffic to 0%
- route all traffic to the stable model
6) Keep at least one last-known-good version ready
Always retain:
- current prod version
- previous prod version
- maybe last 2–3 versions if storage is cheap
Mark versions as:
currentpreviouscandidate
This ensures rollback is not blocked by artifact deletion.
7) Track compatibility between model and features
Many rollbacks fail because the model is fine but the feature pipeline changed.
Version these separately:
- training features
- online feature definitions
- preprocessing code
- tokenizers / vocabularies
- label logic
Store a compatibility matrix:
- model v13 requires feature schema v5
- model v12 works with feature schema v4 or v5
If a rollback model depends on older feature definitions, ensure those still exist.
8) Add monitoring and automated rollback triggers
Monitor:
- prediction latency
- error rate
- CPU/GPU/memory
- input schema violations
- prediction distribution drift
- business metrics
- downstream conversion / fraud loss / CTR
- calibration or confidence collapse
Set rollback triggers such as:
- 5xx rate > X%
- p95 latency > threshold
- prediction distribution deviates significantly
- online metric drops below baseline
- feature missingness spikes
- inference exceptions exceed limit
Example:
- canary receives 5% traffic
- if error rate or business KPI worsens for 10 minutes, automatically rollback to previous version
9) Make rollbacks safe with idempotent deploys
Ensure:
- deploying the same version twice is harmless
- old versions can coexist
- model loading is deterministic
- stateful caches are keyed by version
- requests can be retried safely
Avoid:
- shared mutable global model state
- ad hoc file replacement on running instances
- deleting old versions immediately after deploy
10) Use CI/CD for models, not just code
A model pipeline should include:
- unit tests for preprocessing
- schema tests
- reproducibility checks
- offline evaluation
- bias/fairness checks if relevant
- packaging
- registry registration
- deployment to staging
- canary rollout
- automated monitoring
- rollback on failure
Tools commonly used:
- GitHub Actions / GitLab CI / Jenkins
- MLflow
- Argo CD / Argo Workflows
- Kubeflow Pipelines
- SageMaker Pipelines
- Vertex AI Pipelines
A simple practical architecture
Storage
- Object store for artifacts
- Registry DB for metadata
- Feature store for online/offline features
Serving
- Inference service container loads a model by version
- Config service or environment variable decides active version
- Load balancer / service mesh controls traffic split
Release flow
- Train model
- Store artifact with unique version
- Register model + metrics
- Deploy to staging
- Run tests and validation
- Canary to 5%
- If healthy, promote to 100%
- Keep previous version ready
- On failure, switch alias back or shift traffic back
Recommended minimum setup
If you want the smallest reliable setup, do this:
- Immutable artifacts in S3/GCS
- A registry table with version + stage + metrics
- Production alias pointing to a specific version
- Blue/green or canary rollout
- Automated monitoring and rollback
- Keep last-known-good version always available
Example rollback workflow
- Deploy
model-14to 10% traffic - Monitor for 30 minutes
- Latency and error rate spike
- Automation changes alias:
production -> model-13
- All traffic returns to
model-13 - Investigate
model-14 - Fix issue, retrain, or patch preprocessing
- Re-release as
model-15
Common mistakes to avoid
- overwriting model files in place
- not versioning preprocessing code
- deleting old artifacts too soon
- promoting without canarying
- only watching offline metrics
- not testing schema compatibility
- mixing app deployment version with model version
- no automated rollback path
If you want, I can also give you:
- a reference architecture diagram,
- a database schema for model registry, or
- a step-by-step Kubernetes/MLflow example.