Prompt
why are my end-to-end tests so flaky
Latest observation
Oct 8, 2026
OpenAI APIWeb search: off
End-to-end tests are often flaky because they exercise the whole system at once, so they’re exposed to timing, state, and environment issues that smaller tests avoid. The most common causes are:
- Timing/race conditions: UI or backend isn’t ready when the test checks it.
- Async not properly awaited: promises, network calls, animations, or retries finish at different times.
- Shared state between tests: leftover data, cookies, local storage, database records, or mock state.
- External dependencies: real APIs, third-party services, email, payments, auth, or queues can be slow or inconsistent.
- Environment differences: CI is slower, more parallel, or has different browsers, screen sizes, or resources than local.
- Unstable selectors: using text or DOM structure that changes, instead of stable test IDs.
- Animations and transitions: elements exist but aren’t interactable yet.
- Network variability: requests taking longer, failing temporarily, or returning data in different order.
- Test coupling/order dependence: tests pass only when run in a specific sequence.
- Insufficient isolation: one test affects another through backend state or browser state.
What usually helps:
-
Make tests deterministic
- Stub/mock unstable external calls where possible.
- Seed test data.
- Avoid depending on current time, random IDs, or live external systems.
-
Wait for the right condition, not a fixed delay
- Prefer “element is visible/enabled/request completed” over
sleep(5000).
- Prefer “element is visible/enabled/request completed” over
-
Reset state between tests
- Clear DB, cookies, local storage, caches, queues.
- Create fresh users/accounts per test if needed.
-
Use stable locators
- Prefer
data-testidor equivalent over CSS/XPath based on layout.
- Prefer
-
Reduce UI brittleness
- Disable animations in test mode if possible.
- Avoid asserting too early after navigation or submit.
-
Make CI closer to local or vice versa
- Match browser versions, viewport, and resource limits.
- Run tests in a consistent containerized environment.
-
Split what’s truly E2E from what isn’t
- Use component/integration tests for most behavior.
- Keep E2E for a smaller set of critical flows.
-
Collect evidence
- Capture screenshots, videos, traces, logs, and network events on failure.
- Re-run only the failed test to see if it’s timing-related.
If you want, I can help you diagnose flakiness in your specific setup if you share:
- the test framework (Playwright/Cypress/Selenium/etc.),
- one flaky test,
- and whether it mainly fails locally, in CI, or both.