Prompt
Why are my prompt evals passing but users still complain about bad answers?
Latest observation
Oct 9, 2026
OpenAI APIWeb search: off
A few common reasons:
-
Your evals don’t match real user traffic
- Test prompts are often cleaner, shorter, and more “ideal” than production.
- Users ask vague, messy, multi-intent, or adversarial questions that your eval set may not cover.
-
The evals measure the wrong thing
- You may be scoring for format compliance, keyword overlap, or exact-match style checks, while users care about usefulness, correctness, and completeness.
- A model can “pass” a rubric and still be unhelpful.
-
Your eval set is too small or too narrow
- Passing on 50–200 prompts doesn’t guarantee performance across the long tail.
- Rare but important cases often drive complaints.
-
Synthetic or curated evals are easier than real use
- Human-written benchmark prompts often avoid ambiguity and hidden constraints.
- Real users add context, omit details, or expect the model to infer intent.
-
The system prompt or pipeline behaves differently in production
- Retrieval, tool calls, memory, truncation, routing, guardrails, or latency-based fallbacks can all change behavior.
- The eval may test only the base model, not the full stack.
-
Regressions happen in specific slices
- Overall score can improve while one user segment gets worse.
- Examples: certain languages, domains, device types, or query lengths.
-
Your model is “technically correct” but not practical
- Answers may be accurate but too verbose, too cautious, or poorly structured.
- Users often interpret that as a bad answer.
-
Prompt evals may not capture interaction quality
- Many bad experiences arise over multiple turns: asking clarifying questions, remembering context, recovering from mistakes.
- Single-turn evals miss conversational failure modes.
-
The acceptance threshold is too lenient
- “Pass” may still allow mediocre outputs if the rubric isn’t strict enough.
- Some evals have ceiling effects: everything passes unless it’s obviously broken.
-
User complaints are influenced by expectations
- Even decent outputs can be judged badly if the product promise is unclear or users expect magic.
What to do next
- Compare real user complaints against eval prompts and build a labeled set from them.
- Slice results by:
- intent/domain
- prompt length
- language
- tool usage / retrieval usage
- conversation turn number
- Add evals for:
- ambiguity handling
- refusal behavior
- long-context retention
- tool failure recovery
- follow-up questions
- Measure with a human preference rubric focused on usefulness, correctness, and clarity—not just pass/fail.
- Audit the full production pipeline, not just the model prompt.
If you want, I can help you build a practical debugging checklist or an eval framework that aligns better with user complaints.