Prompt
How do I score summarization quality with human and automated evals?
Latest observation
Scoring summarization quality usually works best if you combine human evaluation and automated evaluation, because they measure different things:
- Humans catch faithfulness, usefulness, and subtle errors.
- Automated metrics give fast, cheap, repeatable signals at scale.
A good setup is to score along a few dimensions rather than one single “summary quality” number.
1) What to evaluate
Common summarization dimensions:
-
Faithfulness / factuality
- Is every claim supported by the source?
- Are there hallucinations, distortions, or omissions of key facts?
-
Coverage / content selection
- Does the summary include the main points?
- Does it miss important information?
-
Conciseness
- Is it short and not repetitive?
- Does it avoid unnecessary detail?
-
Coherence / readability
- Is it well-structured and easy to understand?
- Are pronouns, entities, and transitions clear?
-
Relevance / focus
- Is the summary focused on the source topic and user need?
If you only have room for a few dimensions, prioritize faithfulness and coverage.
2) Human evaluation: how to score it
A. Likert-scale rating
Have annotators rate each summary on a scale, e.g. 1–5:
-
Faithfulness
- 1 = many factual errors
- 3 = minor issues
- 5 = fully supported by source
-
Coverage
- 1 = misses most key points
- 3 = covers some important points
- 5 = covers nearly all important points
-
Coherence
- 1 = hard to read
- 5 = clear and well organized
This is simple and works well if you want aggregate scores.
B. Pairwise preference
Show annotators two summaries for the same source and ask which is better, or whether they are tied.
This often yields more reliable judgments than absolute ratings.
Use it for:
- model comparisons
- product A/B tests
- ranking different prompt/model versions
C. Error annotation
Instead of just rating, ask annotators to mark errors:
- hallucinated fact
- missing key fact
- wrong entity/date/number
- redundant content
- unclear reference
This helps diagnose failures and improve the system.
3) Human eval best practices
- Use a clear rubric. Define each score level with examples.
- Train annotators. Do a calibration round with sample summaries.
- Measure agreement. Use Krippendorff’s alpha, Cohen’s kappa, or pairwise agreement.
- Blind annotators. Hide the system name/model version.
- Randomize order. Avoid position bias in pairwise comparisons.
- Sample representative data. Include easy, hard, and edge-case examples.
- Annotate source access. For faithfulness, annotators should compare summary to source text.
A practical workflow:
- Randomly sample documents.
- Generate summaries from all systems.
- Have 2–3 annotators rate each summary.
- Average scores or use majority vote.
- Report confidence intervals.
4) Automated evaluation: common metrics
A. Lexical overlap metrics
These compare summary words to reference summaries.
- ROUGE-1 / ROUGE-2 / ROUGE-L
- Measures unigram, bigram, and sequence overlap.
- Widely used and easy to compute.
Pros:
- simple
- cheap
- standard baseline
Cons:
- misses paraphrases
- doesn’t measure factuality well
- can reward copying rather than good abstraction
Best for:
- comparison against reference summaries
- quick baseline checks
B. Semantic similarity metrics
These use embeddings or learned models to compare meaning.
- BERTScore
- MoverScore
- other embedding-based metrics
Pros:
- better for paraphrase matching
- more semantic than ROUGE
Cons:
- still not a direct factuality measure
- may correlate imperfectly with human judgment
C. Factuality / consistency metrics
These try to detect whether the summary is supported by the source.
Examples:
- SummaC
- FactCC
- QAFactEval
- BARTScore in some setups
- question-answering based factuality checks
Pros:
- more aligned with hallucination detection
- useful for source-grounded summarization
Cons:
- can be brittle
- often domain-sensitive
- may miss subtle factual errors
D. LLM-based evals
Use a strong language model as a judge with a rubric.
Pros:
- flexible
- can evaluate multiple dimensions
- can give explanations
Cons:
- can be biased
- needs careful prompting and calibration
- should be validated against humans
This is often useful when you need scalable qualitative scoring.
5) A practical scoring framework
A strong setup is:
Human scores
Rate each summary on 1–5 for:
- Faithfulness
- Coverage
- Coherence
- Conciseness
Then compute:
- mean score per dimension
- overall weighted score, e.g.:
- 40% faithfulness
- 30% coverage
- 20% coherence
- 10% conciseness
Automated scores
Track:
- ROUGE for overlap
- BERTScore for semantic similarity
- a factuality metric for hallucination detection
Then correlate these with human scores to see which metrics are meaningful for your task.
6) Example scoring rubric
Here’s a simple faithfulness rubric:
- 5: All statements are supported by the source; no hallucinations.
- 4: Mostly faithful; one minor unsupported detail.
- 3: Some unsupported or distorted content, but main message is correct.
- 2: Several factual errors or misleading statements.
- 1: Mostly hallucinated or contradicted by source.
Coverage rubric:
- 5: Includes all major points.
- 4: Misses one minor point.
- 3: Captures some major points but misses others.
- 2: Captures only a small subset.
- 1: Misses almost all important information.
7) How to combine human and automated evals
A good pattern is:
-
Use automated metrics for iteration
- During development, use ROUGE/BERTScore/factuality checks to quickly compare versions.
-
Use human eval for validation
- Before release, confirm quality with human judgments.
-
Calibrate automated metrics
- Check correlation with human ratings.
- If a metric tracks human faithfulness poorly, don’t rely on it for gating.
-
Use human eval on a held-out benchmark
- Keep a fixed evaluation set to compare future model versions.
8) What not to do
- Don’t rely on ROUGE alone.
- Don’t average everything into one score without checking each dimension.
- Don’t use LLM judges without validating them against humans.
- Don’t evaluate only on easy examples or only on one domain.
- Don’t ignore inter-annotator agreement.
9) Recommended minimal setup
If you want a simple, robust process:
-
Human eval
- 2 annotators
- 1–5 ratings for faithfulness, coverage, coherence
- pairwise preference for model comparisons
-
Automated eval
- ROUGE
- BERTScore
- a factuality metric or LLM judge
-
Report
- mean scores
- agreement
- examples of failure cases
- correlation between automatic and human scores
If you want, I can also give you:
- a ready-to-use annotation rubric,
- a Python evaluation script, or
- a template for an LLM-as-judge prompt.