Prompt
Act as a DevOps architect and recommend a CI/CD setup for a team that deploys to Kubernetes, needs manual approvals for prod, and wants automated rollback on failure.
Latest observation
Here’s a solid CI/CD setup for a team deploying to Kubernetes with manual production approvals and automated rollback on failure.
Recommended approach
Use a GitOps-based deployment model with a CI pipeline for build/test and a CD pipeline for progressive delivery to Kubernetes.
Suggested stack
- CI: GitHub Actions / GitLab CI / Jenkins
- Artifact registry: ECR / GCR / ACR / Harbor
- CD/GitOps: Argo CD
- Progressive delivery / rollback: Argo Rollouts
- Packaging: Helm or Kustomize
- Secrets: External Secrets Operator or sealed secrets
- Policy/quality gates: OPA Gatekeeper / Kyverno, image scanning (Trivy), SAST/DAST as needed
Target flow
1) Developer commits code
- Push to feature branch / PR.
- CI runs:
- lint
- unit tests
- build container image
- security scan
- integration tests if applicable
- On merge to main:
- build immutable image tagged with commit SHA
- push to registry
- update deployment manifest repo or Helm values repo
2) Deploy to lower environments automatically
- Argo CD syncs changes to dev/staging
- Validate:
- smoke tests
- health checks
- API checks
- basic synthetic tests
3) Production requires manual approval
Two common options:
Option A: Manual approval in CI
- After staging passes, pipeline pauses at a manual approval gate
- Approved release updates prod GitOps repo
- Argo CD syncs prod
Option B: Manual promotion via GitOps PR
- Pipeline opens a PR against the prod environment repo
- Human approves and merges
- Argo CD deploys automatically after merge
Recommendation: Option B is usually better for auditability and traceability.
Automated rollback strategy
Use Argo Rollouts with Kubernetes readiness/health checks and progressive traffic shifting.
Rollout pattern
- Deploy new version as canary or blue/green
- Send a small percentage of traffic first
- Measure:
- pod readiness
- error rate
- latency
- custom app metrics from Prometheus
- If metrics fail or rollout analysis fails:
- automatically abort rollout
- rollback to previous stable ReplicaSet
Why this is better than plain Deployment rollbacks
Standard Kubernetes Deployments can rollback, but they do not do intelligent traffic-based progressive delivery. Argo Rollouts gives:
- automated canary analysis
- controlled exposure
- automatic rollback on metric failure
- integration with Prometheus, Datadog, Kayenta, etc.
Recommended pipeline design
CI pipeline stages
- Validate
- format/lint
- unit tests
- Build
- build container image
- tag with commit SHA and semantic version
- Scan
- dependency scan
- image vulnerability scan
- Publish
- push image to registry
- Deploy to dev/staging
- update GitOps repo or Helm chart values
- Argo CD syncs automatically
- Verify
- smoke tests
- integration tests
- Manual approval
- approval gate or PR merge to prod
- Production rollout
- Argo Rollouts canary/blue-green
- Post-deploy monitoring
- alerting and automated rollback on failed analysis
Environment model
Use separate environments:
- dev
- staging
- prod
Recommended best practice:
- separate namespaces at minimum
- ideally separate clusters for prod vs non-prod if compliance/risk warrants it
- separate GitOps overlays per environment
Rollback implementation details
You should define rollback triggers from:
- failed readiness/liveness probes
- failed canary analysis
- increased 5xx rate
- latency thresholds exceeded
- custom business metrics dropping
Example:
- route 10% traffic to new version
- wait 5–10 minutes
- compare error rate and latency
- if metrics exceed thresholds, automatically abort
Also keep:
- last known good image/version
- immutable image tags
- deployment history in Git and cluster controller
Security and governance
Add these controls:
- signed images with cosign
- least-privilege service accounts in Kubernetes
- RBAC for manual approvals
- audit logs for promotion events
- policy checks before prod deploy
- secret scanning in CI
- container vulnerability scanning before publish
High-level architecture
Developer PR
-> CI (test/build/scan)
-> push image to registry
-> update GitOps repo
-> Argo CD deploys to dev/staging
-> tests pass
-> manual approval or PR merge for prod
-> Argo CD syncs prod
-> Argo Rollouts canary/blue-green
-> automated analysis
-> success = promote
-> failure = rollback automatically
My recommendation in one sentence
If you want a modern, reliable setup: GitHub Actions/GitLab CI for CI, Argo CD for GitOps deployments, Argo Rollouts for canary/blue-green with automated rollback, and a manual approval gate via PR merge or pipeline approval before production.
If you want, I can also provide:
- a reference architecture diagram,
- a sample GitHub Actions + Argo CD workflow, or
- a Terraform/Kubernetes implementation blueprint.