Prompt
How do I debug intermittent API test failures?
Latest observation
Oct 6, 2026
OpenAI APIWeb search: off
Intermittent API test failures are usually caused by environment, timing, data, or dependency issues. A good way to debug them is to make the failure observable, repeatable, and isolated.
1. Capture the exact failure
- Log the full request and response:
- method, URL, headers, body
- status code, response body, timing
- Record timestamp, test name, environment, and build/run ID.
- Save any server-side request ID / correlation ID if available.
2. Determine if the failure is reproducible
- Re-run the same test 10–100 times.
- Try it:
- locally
- in CI
- against the same environment
- If it only fails in CI, suspect:
- parallelism
- shared state
- environment differences
- rate limits
- resource contention
3. Check for common intermittent causes
- Race conditions / async timing
- assertions happen before the system is ready
- use polling with timeouts instead of fixed sleeps
- Test data collisions
- two tests use the same user/order/resource
- use unique IDs per test run
- State leakage
- tests depend on leftover data from previous tests
- reset or isolate data between tests
- External dependency instability
- third-party APIs, queues, auth services, databases
- mock or stub dependencies when possible
- Eventual consistency
- API returns success before downstream data is visible
- wait for the condition you need, not just the response
- Flaky network / timeouts
- increase observability around latency and retries
- Rate limiting / throttling
- look for 429s or hidden retry behavior
4. Make the test more deterministic
- Use fixed test fixtures.
- Seed test data explicitly.
- Avoid depending on wall-clock time; inject or mock time if possible.
- Replace arbitrary
sleep(5)with:- polling
- readiness checks
- callbacks/webhooks
- Run tests serially to see whether parallel execution is the cause.
5. Isolate the layer that’s failing
Ask:
- Is the API returning the wrong response?
- Is the backend storing the wrong data?
- Is the test asserting too early?
- Is the client code parsing incorrectly?
Useful tactics:
- Call the API manually with the same payload.
- Re-run against a local/staging environment.
- Compare a passing run vs failing run side by side.
6. Add better instrumentation
- Trace the request through the system.
- Log:
- upstream/downstream calls
- retries
- queue lag
- DB write/read timing
- Add metrics for:
- response latency
- error rates
- retry counts
- queue depth
7. Minimize the test
- Reduce the test to the smallest failing case.
- Remove unrelated steps until only the flake remains.
- A small repro is much easier to diagnose.
8. Use a triage checklist
When a test flakes, check:
- Did it use unique data?
- Did it depend on another test?
- Did it wait for the correct condition?
- Were there network or service errors?
- Did it fail only under parallel load?
- Was there a recent deploy/config change?
9. Handle flakiness carefully
- Don’t just add retries blindly; they can hide real bugs.
- If you retry, do it only as a temporary mitigation while you fix the root cause.
- Track flaky tests separately and make them visible.
10. If you’re using CI
- Compare environment variables, secrets, service versions, and container images.
- Check whether CI runners are overloaded.
- Ensure tests aren’t sharing databases, buckets, queues, or ports.
If you want, I can also give you:
- a step-by-step debugging checklist,
- a template for logging API test failures, or
- examples for pytest/Jest/Postman/Cypress.