Prompt
Kubernetes model serving with rollback
Latest observation
Here’s a practical way to do model serving on Kubernetes with safe rollback.
1) Recommended pattern
Use:
- A versioned model artifact in object storage or a model registry
- A serving container that loads a model based on a version/env var/config
- A Kubernetes Deployment for the inference service
- Progressive delivery for rollout/rollback:
- simple:
kubectl rollout undo - better: Argo Rollouts or Flagger for canary/blue-green
- simple:
2) Basic architecture
- Model registry / storage: S3, GCS, MinIO, MLflow Model Registry, Hugging Face Hub, etc.
- Inference server: FastAPI, TorchServe, TensorFlow Serving, Triton, BentoML, KServe
- Kubernetes:
Deploymentfor app replicasServicefor stable traffic endpointConfigMap/Secretfor model version configReadinessProbeso traffic only goes to healthy pods
3) Simple rollback with Deployment
If each release bakes in or references a model version, Kubernetes can roll back to the previous ReplicaSet.
Example Deployment
apiVersion: apps/v1
kind: Deployment
metadata:
name: model-api
spec:
replicas: 3
revisionHistoryLimit: 5
selector:
matchLabels:
app: model-api
template:
metadata:
labels:
app: model-api
spec:
containers:
- name: model-api
image: myrepo/model-api:1.2.0
env:
- name: MODEL_VERSION
value: "2026-01-01-prod"
ports:
- containerPort: 8000
readinessProbe:
httpGet:
path: /health
port: 8000
initialDelaySeconds: 5
periodSeconds: 5
Update to a new model
kubectl set env deployment/model-api MODEL_VERSION=2026-02-01-prod
kubectl set image deployment/model-api model-api=myrepo/model-api:1.3.0
Roll back
kubectl rollout undo deployment/model-api
Check history:
kubectl rollout history deployment/model-api
4) Better rollback: blue/green
With blue/green, you keep:
- Blue = current production model
- Green = new version
- Switch service selector when green is validated
Pros
- Fast rollback
- No mixed traffic
- Easy validation before switch
Example flow
- Deploy
model-api-v1andmodel-api-v2 - Validate
v2with a separate service or test endpoint - Change production Service selector from
v1tov2 - If issues, switch back immediately
5) Best option for model serving: canary rollout
For ML models, canary is often safest because you can compare:
- latency
- error rate
- business metrics
- prediction drift
- output quality
Tools
- Argo Rollouts
- Flagger
- KServe with canary traffic splitting
Example idea
- Send 5% traffic to new model
- Monitor metrics
- Increase to 25%, 50%, 100%
- Abort and rollback if metrics regress
6) Rollback strategy for model versions
Keep these immutable:
- container image tag
- model artifact version
- config used for preprocessing/postprocessing
A good release should define:
image: myrepo/model-api:1.3.0MODEL_VERSION=2026-02-01FEATURE_SCHEMA_VERSION=4
This makes rollback deterministic.
7) Things to watch for
a) Model compatibility
Ensure old pods can still serve while new pods are warming up.
b) Stateful caches
If you cache embeddings or responses, invalidate carefully.
c) Schema drift
Input schema changes can break rollback if the service expects new fields.
d) Warmup
Large models may need:
- init container
- startup probe
- preloading weights before readiness
Example:
startupProbe:
httpGet:
path: /startup
port: 8000
failureThreshold: 30
periodSeconds: 10
8) KServe example concept
If you use KServe, you can define a model and versioned predictor, then use canary traffic splitting. It’s a strong choice for ML-native Kubernetes serving with rollout support.
9) Minimal recommended production setup
If you want something solid but not overly complex:
- Kubernetes
Deployment ServiceHorizontalPodAutoscalerReadinessProbeandStartupProbe- Versioned model artifacts in S3/MLflow
- GitOps deployment
kubectl rollout undofor basic rollback- Move to Argo Rollouts when you need canary/blue-green
10) Practical advice
If rollback is critical, store:
- the previous model artifact version
- the previous serving image tag
- the previous config map / env values
- validation metrics from the last good release
That way rollback is not just “deploy old code,” but “restore the entire serving state.”
If you want, I can give you:
- a full Kubernetes YAML example
- an Argo Rollouts canary example
- or a KServe deployment with rollback strategy