Prompt
How do I troubleshoot a pipeline that passes tests but fails on deploy?
Latest observation
Here’s a practical way to troubleshoot a pipeline that passes tests but fails during deploy.
1) Compare test environment vs deploy environment
A lot of “works in tests, fails in deploy” issues come from environment differences.
Check:
- Runtime version: Node/Python/Java/.NET version mismatch
- OS/base image: Linux distro, glibc version, shell differences
- Environment variables/secrets: missing or different values
- Dependency versions: lockfile not honored, floating versions, cached packages
- Permissions/IAM: deploy service account can’t access resources
- Network access: deploy target can’t reach registry, cloud APIs, database, etc.
2) Read the deploy logs first
Look for:
- The first real error, not the final cascade of failures
- Any authentication/authorization errors
- Any missing file/config errors
- Any schema/migration errors
- Any timeout or resource limit errors
Tip: If the deploy step is a script, run it with verbose/debug flags:
set -xin shell--debug,--verbose,--trace- increase logging level in the deploy tool
3) Reproduce the deploy step locally or in a similar environment
If tests run in one container and deploy happens in another, try to replicate deploy conditions:
- same container image
- same branch/commit
- same env vars
- same permissions
- same deployment command
If possible, run the deploy command in a disposable environment:
- a CI job
- a staging runner
- a container matching production
4) Check for missing build artifacts
Tests may validate code, but deploy may need files that aren’t created or included.
Verify:
- build output exists
- artifact packaging includes required files
.dockerignore,.gitignore, or packaging config isn’t excluding something important- migrations/config/templates/static assets are present
Examples:
- app compiles, but deploy fails because
dist/wasn’t generated - tests pass, but container image lacks a migration file
- deploy expects a secret/config file that isn’t bundled
5) Look for hidden assumptions in tests
Tests often mock or bypass things deployment actually uses:
- mock database vs real database
- mocked cloud services vs real IAM/resource policies
- local filesystem vs read-only container filesystem
- localhost vs real DNS/TLS/network rules
Ask:
- Are tests using stubs that hide integration issues?
- Do deploy steps run migrations, provisioning, or hooks not covered by tests?
6) Validate configuration and secrets
Deployment failures are often config-related.
Check:
- required env vars are set
- secret names/keys match
- config file format is valid
- YAML/JSON indentation and quoting are correct
- the deploy target has the right region/account/namespace
A common problem is “tests use defaults, deploy uses missing or invalid real values.”
7) Check migration and compatibility issues
If deploy includes database/schema updates:
- Are migrations backward compatible?
- Does app version match schema version?
- Is deploy order correct?
- Are there long-running migrations causing timeouts?
For blue/green or rolling deploys, confirm the old and new versions can coexist.
8) Examine resource limits and timing
Tests may pass quickly, but deploy may fail because:
- startup takes too long
- health checks are too strict
- memory/CPU limits are too low
- container exits before readiness probe passes
- init scripts exceed timeout
Look for:
- OOM kills
- readiness/liveness probe failures
- deployment timeout thresholds
9) Verify access and permissions
If deploy touches cloud resources:
- service principal / IAM role permissions
- Kubernetes RBAC
- registry pull/push permissions
- deployment target write access
- KMS/key vault/secret manager access
If tests run under a different identity than deploy, this is a common root cause.
10) Add targeted checks to the pipeline
If the failure is recurring, add pre-deploy validation:
- config linting
- schema checks
- secret presence checks
- artifact existence checks
- deploy dry-run / plan step
- smoke test after deploy
11) Reduce the problem
If you’re not sure where it breaks:
- split the pipeline into smaller steps
- deploy to staging first
- enable one change at a time
- compare a known-good deploy with the failing one
- bisect recent commits if needed
12) Common patterns to look for
Frequent causes include:
- dependency version drift
- missing env vars/secrets
- IAM/RBAC issues
- packaging mistakes
- migration incompatibility
- differences between test and production configs
- timeout/readiness probe failures
Quick debugging checklist
Use this fast sequence:
- Read deploy logs and find the first error
- Compare test vs deploy environment
- Confirm artifact/package contents
- Validate config/secrets/permissions
- Check migrations and startup/health checks
- Reproduce in a deploy-like environment
If you want, I can also help you build a step-by-step incident checklist for your specific CI/CD system (GitHub Actions, GitLab CI, Jenkins, Azure DevOps, etc.).