Prompt

I'm unhappy with Ragas scores that don't match human judgment

Artificial Intelligence / AI Developer Tools1 observationLast seen Oct 9, 2026

Latest observation

Oct 9, 2026OpenAI APIWeb search: off

That’s a common and valid concern. Ragas scores are useful signals, not ground truth, and they often diverge from human judgment for a few reasons:

Why the mismatch happens

  • Metric mismatch: Ragas may be scoring things like faithfulness, context recall, or answer relevance, while humans may care more about usefulness, completeness, tone, or task success.
  • Weak evaluation context: If the retrieved context is incomplete, a perfectly good answer can score poorly.
  • LLM-as-judge noise: Ragas often relies on LLM judgments, which can be inconsistent, overly sensitive to wording, or biased toward surface features.
  • Domain specificity: In specialized domains, the evaluator model may not understand what “good” looks like.
  • Prompting effects: Small changes in answer phrasing can swing scores without reflecting meaningful quality differences.

What to do about it

  1. Check what each metric is actually measuring

    • Map each Ragas metric to a human-defined quality dimension.
    • If a metric doesn’t align with your goal, don’t use it as a primary KPI.
  2. Build a small human-labeled set

    • Sample 50–200 examples.
    • Have humans rate them on the exact criteria you care about.
    • Compare Ragas against those labels to see where it breaks.
  3. Use Ragas comparatively, not absolutely

    • It’s often better for:
      • comparing prompt versions
      • comparing retrievers
      • regression testing
    • Less reliable as a final “quality score.”
  4. Tune thresholds per task

    • A 0.8 in one domain may be mediocre, while in another it may be excellent.
    • Calibrate against your own human ratings.
  5. Add task-specific evaluation

    • For example:
      • exact answer correctness
      • citation accuracy
      • tool success
      • business rule compliance
      • refusal quality
    • These often correlate better with human judgment than generic metrics.
  6. Inspect failure cases

    • Look at examples where Ragas and humans disagree.
    • Patterns usually emerge: missing nuance, overly harsh faithfulness scoring, wrong retrieval assumptions, etc.

Practical recommendation

Use a hybrid eval stack:

  • Ragas for automated, scalable regression checks
  • Human evals for calibration and periodic audits
  • Task-specific rules/metrics for what matters operationally

If you want, I can help you:

  • diagnose which Ragas metric is drifting from human judgment,
  • design a human evaluation rubric,
  • or create a calibration plan to make scores more aligned.

Brands

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.