Prompt
How do I evaluate summarization quality without exact match scores?
Latest observation
Evaluating summarization quality without exact-match scores usually means using a mix of automatic metrics, semantic similarity measures, and human evaluation. Exact match is rarely appropriate for summarization because many different summaries can be valid.
1. Use overlap-based automatic metrics
These compare generated summaries to one or more reference summaries, but allow partial matches:
- ROUGE-N: n-gram overlap, especially ROUGE-1 and ROUGE-2
- ROUGE-L: longest common subsequence overlap
- ROUGE-S: skip-bigram overlap
These are the most common baseline metrics for summarization, though they mostly measure lexical similarity rather than meaning.
2. Use semantic similarity metrics
If you want to go beyond word overlap, use metrics that compare meaning:
- BERTScore: compares token embeddings between candidate and reference summaries
- MoverScore: uses contextual embeddings and optimal transport
- BLEURT: learned metric trained to predict human judgments
- BARTScore: scores summaries using a seq2seq model’s likelihood
These tend to better reflect paraphrases and semantic equivalence.
3. Evaluate factual consistency / faithfulness
A summary can be fluent and similar to the reference but still be wrong. To check factual quality:
- QAGS / QAEval: generate questions from the summary and answer them using source and summary
- FactCC: checks whether summary statements are supported by the source
- SummaC: consistency checking with entailment-style models
- Entity-level checks: verify names, dates, numbers, and relations against the source
This is especially important for abstractive summarization.
4. Human evaluation
Human judges can rate summaries on dimensions such as:
- Coherence: is the summary well-structured?
- Informativeness / coverage: does it capture the important points?
- Faithfulness / factual accuracy: is it supported by the source?
- Fluency: is it grammatical and readable?
- Conciseness: does it avoid unnecessary details?
Common formats:
- Likert scale ratings
- Pairwise preference judgments
- Best-worst scaling
Human evaluation is the most reliable, but it’s expensive and slower.
5. Task-based evaluation
If summaries are used downstream, measure utility directly:
- Can users answer questions about the document from the summary?
- Does the summary help with retrieval or decision-making?
- Does it improve speed/accuracy on a downstream task?
This is often the most meaningful evaluation when summaries are used in an application.
6. Use multiple references if possible
Summarization has many valid outputs. If you have only one reference, automatic metrics can be overly harsh. Multiple reference summaries improve evaluation robustness.
Practical recommendation
A solid evaluation setup is:
- ROUGE for standard comparison
- BERTScore or BLEURT for semantic similarity
- Faithfulness metric like SummaC or FactCC
- Human evaluation on a small sample
If you want, I can also help you design an evaluation pipeline for a specific summarization project, such as news, meetings, or long-document summarization.