Prompt

Why are my prompt evals passing but users still complain about bad answers?

Technology / Observability1 observationLast seen Oct 9, 2026

Latest observation

Oct 9, 2026OpenAI APIWeb search: off

A few common reasons:

  1. Your evals don’t match real user traffic

    • Test prompts are often cleaner, shorter, and more “ideal” than production.
    • Users ask vague, messy, multi-intent, or adversarial questions that your eval set may not cover.
  2. The evals measure the wrong thing

    • You may be scoring for format compliance, keyword overlap, or exact-match style checks, while users care about usefulness, correctness, and completeness.
    • A model can “pass” a rubric and still be unhelpful.
  3. Your eval set is too small or too narrow

    • Passing on 50–200 prompts doesn’t guarantee performance across the long tail.
    • Rare but important cases often drive complaints.
  4. Synthetic or curated evals are easier than real use

    • Human-written benchmark prompts often avoid ambiguity and hidden constraints.
    • Real users add context, omit details, or expect the model to infer intent.
  5. The system prompt or pipeline behaves differently in production

    • Retrieval, tool calls, memory, truncation, routing, guardrails, or latency-based fallbacks can all change behavior.
    • The eval may test only the base model, not the full stack.
  6. Regressions happen in specific slices

    • Overall score can improve while one user segment gets worse.
    • Examples: certain languages, domains, device types, or query lengths.
  7. Your model is “technically correct” but not practical

    • Answers may be accurate but too verbose, too cautious, or poorly structured.
    • Users often interpret that as a bad answer.
  8. Prompt evals may not capture interaction quality

    • Many bad experiences arise over multiple turns: asking clarifying questions, remembering context, recovering from mistakes.
    • Single-turn evals miss conversational failure modes.
  9. The acceptance threshold is too lenient

    • “Pass” may still allow mediocre outputs if the rubric isn’t strict enough.
    • Some evals have ceiling effects: everything passes unless it’s obviously broken.
  10. User complaints are influenced by expectations

    • Even decent outputs can be judged badly if the product promise is unclear or users expect magic.

What to do next

  • Compare real user complaints against eval prompts and build a labeled set from them.
  • Slice results by:
    • intent/domain
    • prompt length
    • language
    • tool usage / retrieval usage
    • conversation turn number
  • Add evals for:
    • ambiguity handling
    • refusal behavior
    • long-context retention
    • tool failure recovery
    • follow-up questions
  • Measure with a human preference rubric focused on usefulness, correctness, and clarity—not just pass/fail.
  • Audit the full production pipeline, not just the model prompt.

If you want, I can help you build a practical debugging checklist or an eval framework that aligns better with user complaints.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.