Prompt

How do I check whether a judge model agrees with human ratings?

Artificial Intelligence / AI Developer Tools1 observationLast seen Oct 10, 2026

Latest observation

Oct 10, 2026OpenAI APIWeb search: off

To check whether a judge model agrees with human ratings, you usually compare the model’s scores/rankings against human annotations with an agreement metric.

1) First decide what kind of outputs you have

  • Both human and judge give categorical labels (e.g. good/bad, 1–5 stars)
  • Both give rankings or pairwise preferences (A better than B)
  • Both give numeric scores (e.g. 0–10)

The best agreement measure depends on this.


2) Common ways to measure agreement

If labels are categorical

Use:

  • Accuracy: percent of exact matches
  • Cohen’s kappa: agreement beyond chance for 2 raters
  • Fleiss’ kappa: if multiple human raters
  • Weighted kappa: if labels are ordered, like 1–5 stars
  • Krippendorff’s alpha: good for multiple raters, missing data

If outputs are numeric

Use:

  • Pearson correlation: linear relationship
  • Spearman correlation: rank-order agreement
  • Mean absolute error (MAE) or RMSE: score closeness

If outputs are pairwise preferences

Use:

  • Pairwise agreement rate
  • Win rate: how often the judge picks the same option as humans
  • Kendall’s tau or Spearman if converted to rankings

3) A practical evaluation workflow

  1. Collect a dataset where:
    • each example has a human rating
    • the judge model gives a rating on the same example
  2. Compare them with an appropriate metric
  3. Check agreement overall and by subgroup/task type
  4. Inspect disagreements manually to see systematic bias
  5. If possible, compare judge model to multiple humans, not just one

4) Important caveat

A judge model can have high correlation with human ratings but still be poorly calibrated.
For example, it may rank examples similarly to humans but consistently rate everything too high.

So it’s useful to report both:

  • agreement metrics like kappa/correlation
  • calibration/error metrics like MAE or confusion matrix

5) Example interpretations

  • Accuracy = 85%: exact label match on 85% of items
  • Cohen’s kappa = 0.72: substantial agreement beyond chance
  • Spearman = 0.81: judge and humans rank items similarly
  • MAE = 0.6 on a 1–5 scale: average prediction is off by 0.6 stars

6) If you want a simple recommendation

  • Binary labels: use accuracy + Cohen’s kappa
  • 1–5 ratings: use weighted kappa + Spearman + MAE
  • Multiple human raters: use Krippendorff’s alpha or majority vote plus agreement metrics
  • Pairwise judgments: use pairwise win rate + Kendall/Spearman

If you want, I can also give you:

  • a Python snippet to compute these metrics, or
  • advice on which metric to use for your exact rating setup.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.