Prompt
Kubernetes rollbacks and canary deployments best practices
Latest observation
Oct 6, 2026
OpenAI APIWeb search: off
Here are practical best practices for Kubernetes rollbacks and canary deployments.
1) Rollback best practices
Use declarative, versioned deployments
- Keep manifests in Git and deploy from a known revision.
- Prefer GitOps or CI/CD pipelines so every release is traceable.
- Tag container images immutably (
myapp:1.2.3, or better, image digest).
Make rollbacks fast and safe
- Roll back by changing the desired state, not manually editing live resources.
- For Deployments, use:
kubectl rollout undo deployment/<name>- or re-apply the previous manifest/image tag.
Keep old ReplicaSets around briefly
- Kubernetes Deployments retain previous ReplicaSets by default, which makes rollback easier.
- Avoid aggressive cleanup of old revisions until the new version is proven healthy.
Set good health probes
- readinessProbe: remove bad pods from service traffic quickly.
- livenessProbe: restart stuck pods.
- startupProbe: avoid premature restarts for slow-starting apps.
- Bad probes often make rollbacks slower or less reliable.
Use rollout controls
- Set
maxUnavailableandmaxSurgecarefully for Deployments. - Common safe default:
maxUnavailable: 0maxSurge: 1
- This reduces impact during updates and makes rollback behavior more predictable.
Automate rollback triggers
Rollback automatically when:
- error rate rises
- latency exceeds SLOs
- readiness drops
- a canary metric fails
- synthetic checks fail
Test rollback before production
- Practice rollback in staging.
- Verify:
- old image still exists
- schema changes are backward compatible
- app tolerates mixed-version traffic
Be careful with database migrations
- Avoid non-backward-compatible schema changes in the same release as app changes.
- Use expand/contract patterns:
- add new columns/fields
- deploy new app
- migrate data
- remove old fields later
2) Canary deployment best practices
Start with a small traffic percentage
- Begin at 1–5% traffic.
- Increase gradually if metrics stay healthy.
Use objective success criteria
Define clear metrics before rollout:
- error rate
- p95/p99 latency
- CPU/memory anomalies
- business KPIs
- request success by endpoint
Compare canary vs baseline
- Always compare the canary to a stable control version.
- Metrics should be segmented by version, pod label, or route.
Route traffic carefully
Common approaches:
- Ingress/controller-based routing
- Service mesh (Istio, Linkerd, etc.)
- Argo Rollouts / Flagger
- Weighted load balancer or HTTP routing
Keep canaries statistically meaningful
- Ensure enough traffic and duration to detect issues.
- Don’t promote a canary too quickly if traffic is low or patterns are uneven.
Limit blast radius
- Canary only one region/namespace/customer segment if possible.
- Avoid spreading the canary across too many dimensions at once.
Make pods comparable
- Same resource limits/requests
- Same node affinity/tolerations if possible
- Same autoscaling settings
- Similar configuration except for the version under test
Watch for hidden regressions
Canary failures may show up in:
- cache behavior
- memory leaks
- slow shutdowns
- connection pool exhaustion
- background jobs
- rare code paths
Automate promotion and rollback
Use progressive delivery tooling to:
- pause
- analyze metrics
- promote
- abort and rollback automatically
3) Recommended deployment patterns
Blue/Green
- Two full environments: one live, one idle.
- Pros: simple rollback, instant switch.
- Cons: more infrastructure cost.
Canary
- Small percentage of traffic to new version.
- Pros: safest for catching regressions.
- Cons: more complex routing/observability.
Rolling update
- Default Kubernetes approach.
- Pros: simple.
- Cons: less control than canary; rollback is possible but traffic shaping is limited.
4) Tools commonly used
- Native Kubernetes Deployment: basic rolling updates and rollbacks
- Argo Rollouts: advanced canary/blue-green with analysis
- Flagger: automated progressive delivery
- Istio / Linkerd: traffic splitting and telemetry
- Helm/Kustomize + GitOps (Argo CD / Flux): release management
5) Practical checklist
Before release:
- Image is immutable and versioned
- Health probes are correct
- Metrics/alerts are in place
- Rollback path is tested
- DB migrations are backward compatible
- Canary plan and thresholds are defined
During release:
- Send small percentage to canary
- Monitor error rate, latency, saturation
- Compare against baseline
- Pause if anything looks off
After release:
- Promote only after sustained success
- Keep old version available until stable
- Postmortem any failed rollout
6) Common pitfalls
- Rolling out app + breaking DB migration together
- No readiness probe, so bad pods receive traffic
- Relying only on pod status, not application metrics
- Using mutable tags like
latest - Promoting canaries based on time alone, not metrics
- Cleaning up old revisions too early
- Forgetting background jobs, cron jobs, and consumers in canary scope
If you want, I can also provide:
- a Kubernetes Deployment YAML with rollback-friendly settings,
- an Argo Rollouts canary example, or
- a production release checklist template.