Prompt
I'm unhappy with Ragas scores that don't match human judgment
Latest observation
Oct 9, 2026
OpenAI APIWeb search: off
That’s a common and valid concern. Ragas scores are useful signals, not ground truth, and they often diverge from human judgment for a few reasons:
Why the mismatch happens
- Metric mismatch: Ragas may be scoring things like faithfulness, context recall, or answer relevance, while humans may care more about usefulness, completeness, tone, or task success.
- Weak evaluation context: If the retrieved context is incomplete, a perfectly good answer can score poorly.
- LLM-as-judge noise: Ragas often relies on LLM judgments, which can be inconsistent, overly sensitive to wording, or biased toward surface features.
- Domain specificity: In specialized domains, the evaluator model may not understand what “good” looks like.
- Prompting effects: Small changes in answer phrasing can swing scores without reflecting meaningful quality differences.
What to do about it
-
Check what each metric is actually measuring
- Map each Ragas metric to a human-defined quality dimension.
- If a metric doesn’t align with your goal, don’t use it as a primary KPI.
-
Build a small human-labeled set
- Sample 50–200 examples.
- Have humans rate them on the exact criteria you care about.
- Compare Ragas against those labels to see where it breaks.
-
Use Ragas comparatively, not absolutely
- It’s often better for:
- comparing prompt versions
- comparing retrievers
- regression testing
- Less reliable as a final “quality score.”
- It’s often better for:
-
Tune thresholds per task
- A 0.8 in one domain may be mediocre, while in another it may be excellent.
- Calibrate against your own human ratings.
-
Add task-specific evaluation
- For example:
- exact answer correctness
- citation accuracy
- tool success
- business rule compliance
- refusal quality
- These often correlate better with human judgment than generic metrics.
- For example:
-
Inspect failure cases
- Look at examples where Ragas and humans disagree.
- Patterns usually emerge: missing nuance, overly harsh faithfulness scoring, wrong retrieval assumptions, etc.
Practical recommendation
Use a hybrid eval stack:
- Ragas for automated, scalable regression checks
- Human evals for calibration and periodic audits
- Task-specific rules/metrics for what matters operationally
If you want, I can help you:
- diagnose which Ragas metric is drifting from human judgment,
- design a human evaluation rubric,
- or create a calibration plan to make scores more aligned.
Brands