Prompt
How do I version and roll back model deployments?
Latest observation
A good way to version and roll back model deployments is to treat each deployed model like a releasable artifact with a unique, immutable version, and deploy through an indirection layer so you can switch versions quickly.
1) Version the right things
Track versions for:
- Model artifact: weights, checkpoint, ONNX/TensorRT file, etc.
- Training code: git commit SHA
- Training data: dataset version or snapshot ID
- Feature schema: input column names/types, preprocessing logic
- Environment: container image or dependency lockfile
- Config: thresholds, prompts, decoding settings, etc.
A common release identifier might look like:
model_name:1.8.3
or
model_name@gitsha+data2026_01_15+imgsha
Best practice: make deployed artifacts immutable. Never overwrite version 1.8.3; publish 1.8.4 instead.
2) Use a model registry
Store versions in a registry such as:
- MLflow Model Registry
- SageMaker Model Registry
- Vertex AI Model Registry
- Azure ML registry
- Custom registry in S3/GCS + metadata DB
For each model version, record:
- artifact URI
- metrics
- validation status
- approved/staged/production state
- compatibility notes
- rollback candidate
3) Deploy with traffic indirection
Don’t point clients directly at a specific file. Put a routing layer in front:
- service name:
fraud-model-prod - behind it: version
1.8.3,1.8.4, etc.
Deployment patterns:
- Blue/green: keep old version live until new one is healthy
- Canary: send a small percentage to new version first
- Shadow: run new version in parallel without affecting responses
This makes rollback trivial: just re-point traffic to the previous version.
4) Keep old versions available
Rollback only works if you retain:
- the previous model artifact
- its exact container/runtime
- its feature preprocessing code
- any config needed to reproduce behavior
Set retention policies carefully:
- keep production and last N versions
- archive older versions if storage is expensive
- never delete the last known-good version without a replacement
5) Define rollback triggers
Roll back automatically or manually when you see:
- latency regressions
- increased error rates
- degraded accuracy/business metrics
- input drift or data quality failures
- memory/CPU/GPU issues
- unexpected output distribution changes
Good systems use:
- health checks
- online metrics monitoring
- alerting
- SLO-based rollback criteria
6) Rollback procedure
A simple rollback flow:
- Detect issue in version
vNew - Verify last known good version
vOld - Switch traffic from
vNewback tovOld - Freeze
vNewfor investigation - Postmortem and fix forward
- Re-release as
vNew+1
If using Kubernetes:
- keep separate deployments per version, or
- use a stable service selector / ingress route you can update quickly
If using a model server:
- maintain aliases like
prod -> 1.8.3 - rollback by changing alias back to
1.8.2
7) Test compatibility before promoting
Before a model is eligible for production:
- unit test preprocessing and postprocessing
- validate feature schema
- run offline evaluation against holdout data
- run canary or shadow deployment
- verify backward compatibility with client expectations
8) Practical naming scheme
Example:
model_namemodel_name:1.2.0semantic version- metadata:
git_sha=a1b2c3ddata_version=2026-01-01feature_spec=v7container=sha256:...
Use semantic versioning:
- major: breaking changes
- minor: improved performance, compatible
- patch: bug fixes, non-functional changes
9) Minimal rollback design
If you want the simplest robust setup:
- Registry stores versions
v1,v2,v3 - Production points to an alias
current - Promote new version by changing
current - Roll back by repointing
currentto prior version - Never delete prior artifacts immediately
10) Recommended operational checklist
- Immutable versioned artifacts
- Registry with metadata
- Production alias or routing layer
- Canary/shadow deployment
- Health + quality monitoring
- One-click rollback to last known good version
- Retain reproducibility for old versions
If you want, I can also show:
- a Kubernetes-based rollout/rollback setup
- an MLflow/SageMaker/Vertex AI versioning pattern
- or a sample CI/CD pipeline for model deployments.