Prompt

How do I fix API tests that pass locally but fail in CI?

Technology / Testing & QA1 observationLast seen Oct 6, 2026

Latest observation

Oct 6, 2026OpenAI APIWeb search: off

When API tests pass locally but fail in CI, the cause is usually environment differences, timing/race issues, or shared state. Here’s a practical checklist to find and fix them.

1) Compare local vs CI environments

Look for differences in:

  • Node/Python/Java version
  • OS / container image
  • Environment variables
  • Dependency versions
  • Database/cache versions
  • Locale/timezone
  • Network access / DNS / proxies

Fix: Pin versions and make CI match local as closely as possible.

  • Use a lockfile (package-lock.json, poetry.lock, requirements.txt with pinned versions, etc.)
  • Use the same runtime version in CI as local
  • Prefer Docker for consistent test environments

2) Eliminate flaky timing issues

CI is often slower, so tests that “just barely work” locally fail there.

Symptoms:

  • Assertions happen before async work finishes
  • Polling/retries are too short
  • Sleep-based tests are unreliable

Fix:

  • Await all async operations
  • Replace sleep(1) with condition-based waits
  • Increase timeouts if necessary, but don’t hide bugs
  • Wait for explicit signals: response status, DB record exists, job completed, etc.

Bad:

await sleep(1000);
expect(result.status).toBe('done');

Better:

await waitFor(() => expect(result.status).toBe('done'));

3) Remove test interdependence

Tests may pass alone but fail in CI when run:

  • in parallel
  • in a different order
  • after another test changed shared data

Common causes:

  • Reusing the same user/email/resource names
  • Not cleaning up test data
  • Global mutable state
  • Assuming DB starts empty

Fix:

  • Make each test self-contained
  • Generate unique test data
  • Clean up after tests
  • Reset DB between tests if needed
  • Avoid reliance on order

4) Check for parallel execution problems

CI often runs tests in parallel, exposing race conditions.

Examples:

  • Two tests insert the same record
  • Shared temp files
  • Shared ports
  • Global mocks not restored
  • Same external service account used by many tests

Fix:

  • Make resources unique per test
  • Use isolated databases/schemas
  • Serialize tests that truly can’t run concurrently
  • Restore mocks after each test

5) Verify external dependencies

If your API tests call:

  • real third-party APIs
  • real S3/Redis/Kafka
  • real auth services

then CI failures may be due to:

  • rate limiting
  • network instability
  • missing credentials
  • sandbox vs production differences

Fix:

  • Mock external services in unit/integration tests
  • Use dedicated test environments for true end-to-end tests
  • Add retries only where appropriate
  • Ensure CI secrets and endpoints are correctly configured

6) Make assertions more robust

Tests can be too strict or depend on unstable details.

Examples:

  • Comparing full JSON payloads when some fields are dynamic
  • Expecting exact timestamps, IDs, ordering, or formatting
  • Depending on map/object key order

Fix:

  • Assert only what matters
  • Normalize dynamic fields
  • Use partial matches when appropriate

Example:

expect(response.body).toMatchObject({
  status: 'ok',
  user: { email: 'test@example.com' }
});

7) Capture better CI diagnostics

If CI just says “failed,” make it easier to debug.

Add:

  • Full request/response logs
  • Server logs
  • DB state on failure
  • Screenshots only if relevant, but for APIs use payload dumps
  • Artifacts from CI for failed runs

Fix:

  • Print the response body and status on failure
  • Save logs as artifacts
  • Run failed tests in verbose mode

8) Reproduce CI locally

The fastest path to fixing is often to run the tests in the same way CI does.

Try:

  • Run tests in Docker using the same image as CI
  • Use the same env vars
  • Run the exact test command from CI
  • Run tests in parallel locally if CI does
  • Run the whole suite, not just the single passing test

9) Look for hidden assumptions

Common hidden assumptions:

  • Current time/date
  • Current locale/timezone
  • Current working directory
  • Filesystem case sensitivity
  • Availability of certain ports
  • A “fresh” database or cache

Fix:

  • Freeze time in tests
  • Set timezone explicitly
  • Use relative paths carefully
  • Don’t assume default state

10) Use retries carefully

Retries can reduce noise for flaky integration tests, but they’re not a real fix.

Use retries only if:

  • the operation is genuinely eventually consistent
  • external services are occasionally unstable

Avoid retries for:

  • deterministic test failures
  • logic bugs
  • race conditions you can properly synchronize

A good debugging workflow

  1. Identify whether it’s a timing, data, or environment issue
  2. Run tests in CI-like conditions locally
  3. Check logs and failing response payloads
  4. Isolate the failing test
  5. Remove shared state and race conditions
  6. Pin versions and dependencies
  7. Mock unstable external dependencies
  8. Re-run the full suite repeatedly

Quick checklist

  • Same runtime version locally and in CI
  • Same dependency lockfile used
  • Tests don’t rely on order
  • Unique test data for every run
  • No fixed sleeps for async work
  • External APIs are mocked or controlled
  • Assertions ignore dynamic fields where appropriate
  • CI logs enough detail to debug failures

If you want, I can help you debug a specific failing test or CI log and point out the likely root cause.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.