Prompt
How do I check whether a judge model agrees with human ratings?
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
To check whether a judge model agrees with human ratings, you usually compare the model’s scores/rankings against human annotations with an agreement metric.
1) First decide what kind of outputs you have
- Both human and judge give categorical labels (e.g. good/bad, 1–5 stars)
- Both give rankings or pairwise preferences (A better than B)
- Both give numeric scores (e.g. 0–10)
The best agreement measure depends on this.
2) Common ways to measure agreement
If labels are categorical
Use:
- Accuracy: percent of exact matches
- Cohen’s kappa: agreement beyond chance for 2 raters
- Fleiss’ kappa: if multiple human raters
- Weighted kappa: if labels are ordered, like 1–5 stars
- Krippendorff’s alpha: good for multiple raters, missing data
If outputs are numeric
Use:
- Pearson correlation: linear relationship
- Spearman correlation: rank-order agreement
- Mean absolute error (MAE) or RMSE: score closeness
If outputs are pairwise preferences
Use:
- Pairwise agreement rate
- Win rate: how often the judge picks the same option as humans
- Kendall’s tau or Spearman if converted to rankings
3) A practical evaluation workflow
- Collect a dataset where:
- each example has a human rating
- the judge model gives a rating on the same example
- Compare them with an appropriate metric
- Check agreement overall and by subgroup/task type
- Inspect disagreements manually to see systematic bias
- If possible, compare judge model to multiple humans, not just one
4) Important caveat
A judge model can have high correlation with human ratings but still be poorly calibrated.
For example, it may rank examples similarly to humans but consistently rate everything too high.
So it’s useful to report both:
- agreement metrics like kappa/correlation
- calibration/error metrics like MAE or confusion matrix
5) Example interpretations
- Accuracy = 85%: exact label match on 85% of items
- Cohen’s kappa = 0.72: substantial agreement beyond chance
- Spearman = 0.81: judge and humans rank items similarly
- MAE = 0.6 on a 1–5 scale: average prediction is off by 0.6 stars
6) If you want a simple recommendation
- Binary labels: use accuracy + Cohen’s kappa
- 1–5 ratings: use weighted kappa + Spearman + MAE
- Multiple human raters: use Krippendorff’s alpha or majority vote plus agreement metrics
- Pairwise judgments: use pairwise win rate + Kendall/Spearman
If you want, I can also give you:
- a Python snippet to compute these metrics, or
- advice on which metric to use for your exact rating setup.