Prompt
Can you suggest a way to evaluate LLM summaries against references and human judgments?
Latest observation
Oct 9, 2026
OpenAI APIWeb search: off
Yes — a solid way is to use a multi-level evaluation setup that combines:
- Automatic overlap/semantic metrics against references
- Human judgment on a small but representative subset
- Correlation analysis to see which automatic metrics best track human preferences
A practical approach
1) Compare summaries to references with several metrics
Use a mix of metrics, since each captures something different:
- ROUGE-1/2/L: good for surface overlap
- BERTScore: captures semantic similarity better than ROUGE
- BLEURT or COMET: often better aligned with human judgments, especially for generation quality
- QAEval / FactCC / SummaC: useful for factual consistency, which is often what humans care about most in summaries
2) Collect human judgments
Ask annotators to rate summaries on dimensions such as:
- Faithfulness / factual consistency
- Coverage / informativeness
- Coherence / readability
- Conciseness
- Optional: overall preference
Use a simple rubric, for example a 1–5 scale, or pairwise comparisons between two summaries.
3) Measure agreement with human judgments
To see whether the automatic metrics are meaningful:
- Compute Spearman or Kendall correlation between metric scores and human scores
- If using pairwise preferences, check pairwise accuracy of the metric
- If multiple annotators, report inter-annotator agreement too
Recommended evaluation design
If you want one strong setup, do this:
- Use reference summaries and score with ROUGE + BERTScore + a factuality metric
- Get human ratings for ~100–300 examples
- Compare metric rankings to human rankings using Kendall/Spearman
- Pick the metric that correlates best with human judgments for future large-scale evaluation
Important caveat
Reference-based metrics can miss good summaries that express the same content differently, so human evaluation is still important. For summarization, factuality often matters more than lexical similarity.
If you want, I can also suggest:
- a full evaluation protocol,
- a Python implementation for computing these metrics,
- or a human annotation rubric.