Prompt
judge model human agreement llm
Latest observation
Oct 9, 2026
OpenAI APIWeb search: off
If you want to judge model–human agreement for an LLM, the usual approach is:
1) Define the task
You need to know what kind of agreement you mean:
- Classification: do model and humans assign the same label?
- Rating: do they give similar scores?
- Preference / ranking: do they choose the same option?
- Free-text evaluation: do they produce comparable judgments or annotations?
2) Use an agreement metric
Common choices:
-
Accuracy / exact match
Simple percentage of identical outputs. -
Cohen’s kappa
Good for 1 human vs 1 model on categorical labels, adjusts for chance agreement. -
Fleiss’ kappa
For multiple human annotators plus the model. -
Krippendorff’s alpha
Flexible, works with missing data and different data types. -
Pearson / Spearman correlation
For continuous or ordinal ratings. -
MAE / RMSE
For numeric scores. -
Pairwise agreement / win rate
For preference judgments.
3) Compare against human baseline
Don’t just compare model vs human once. Also measure:
- Human–human agreement
- Model–human agreement
- Whether model agreement is within the range of human annotators
4) Check statistical significance
Use:
- Bootstrap confidence intervals
- Permutation tests
- McNemar’s test for paired classification comparisons
5) Inspect errors
Agreement metrics can hide systematic issues:
- Bias toward one label
- Overuse of neutral responses
- Better performance on easy cases only
6) If using an LLM as a judge
Make sure to:
- Fix prompts
- Use temperature 0 for deterministic evaluation
- Randomize order of options to avoid position bias
- Evaluate on blinded examples
- Validate against a gold standard or expert annotation
If you want, I can also give you:
- a recommended metric for your exact task, or
- a Python evaluation script for model–human agreement.