Prompt

why are my end-to-end tests so flaky

Technology / Testing & QA1 observationLast seen Oct 8, 2026

Latest observation

Oct 8, 2026OpenAI APIWeb search: off

End-to-end tests are often flaky because they exercise the whole system at once, so they’re exposed to timing, state, and environment issues that smaller tests avoid. The most common causes are:

  • Timing/race conditions: UI or backend isn’t ready when the test checks it.
  • Async not properly awaited: promises, network calls, animations, or retries finish at different times.
  • Shared state between tests: leftover data, cookies, local storage, database records, or mock state.
  • External dependencies: real APIs, third-party services, email, payments, auth, or queues can be slow or inconsistent.
  • Environment differences: CI is slower, more parallel, or has different browsers, screen sizes, or resources than local.
  • Unstable selectors: using text or DOM structure that changes, instead of stable test IDs.
  • Animations and transitions: elements exist but aren’t interactable yet.
  • Network variability: requests taking longer, failing temporarily, or returning data in different order.
  • Test coupling/order dependence: tests pass only when run in a specific sequence.
  • Insufficient isolation: one test affects another through backend state or browser state.

What usually helps:

  1. Make tests deterministic

    • Stub/mock unstable external calls where possible.
    • Seed test data.
    • Avoid depending on current time, random IDs, or live external systems.
  2. Wait for the right condition, not a fixed delay

    • Prefer “element is visible/enabled/request completed” over sleep(5000).
  3. Reset state between tests

    • Clear DB, cookies, local storage, caches, queues.
    • Create fresh users/accounts per test if needed.
  4. Use stable locators

    • Prefer data-testid or equivalent over CSS/XPath based on layout.
  5. Reduce UI brittleness

    • Disable animations in test mode if possible.
    • Avoid asserting too early after navigation or submit.
  6. Make CI closer to local or vice versa

    • Match browser versions, viewport, and resource limits.
    • Run tests in a consistent containerized environment.
  7. Split what’s truly E2E from what isn’t

    • Use component/integration tests for most behavior.
    • Keep E2E for a smaller set of critical flows.
  8. Collect evidence

    • Capture screenshots, videos, traces, logs, and network events on failure.
    • Re-run only the failed test to see if it’s timing-related.

If you want, I can help you diagnose flakiness in your specific setup if you share:

  • the test framework (Playwright/Cypress/Selenium/etc.),
  • one flaky test,
  • and whether it mainly fails locally, in CI, or both.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.