Prompt
I'm building a summarization system and need to score output quality
Latest observation
To score summarization quality, you’ll usually want a mix of automatic metrics and human evaluation.
1) Automatic metrics
Reference-based metrics
Used when you have one or more gold summaries.
-
ROUGE-1 / ROUGE-2 / ROUGE-L
- Measures overlap of unigrams, bigrams, and longest common subsequence.
- Common for extractive and abstractive summarization.
- Good baseline, but can miss paraphrases and semantic equivalence.
-
BERTScore
- Compares contextual embeddings between candidate and reference.
- Better at capturing semantic similarity than ROUGE.
-
MoverScore / BLEURT / BARTScore
- More semantic-aware metrics.
- Often correlate better with human judgments than plain overlap metrics.
Reference-free metrics
Useful when you don’t have gold summaries.
-
Factual consistency / hallucination detectors
- Check whether summary claims are supported by source text.
- Examples: entailment-based scores, QA-based evaluation, NLI models.
-
Compression / coverage heuristics
- Measures how much source content is represented without being too long.
- Can be useful as diagnostics, not final quality scores.
2) Human evaluation dimensions
Common criteria:
- Faithfulness / factuality
- Is the summary consistent with the source?
- Coverage / informativeness
- Does it include the important points?
- Conciseness
- Is it brief without unnecessary detail?
- Readability / fluency
- Is it grammatical and coherent?
- Overall usefulness
- Would a person consider it a good summary?
A typical rubric is 1–5 for each dimension, then combine into an overall score.
3) Practical recommendation
For a production scoring system, I’d suggest:
- ROUGE for a quick baseline
- BERTScore or BLEURT for semantic similarity
- Factual consistency check against the source
- Human review for a sampled subset
A simple composite score might look like:
[ \text{Score} = 0.4 \cdot \text{Faithfulness} + 0.3 \cdot \text{Coverage} + 0.2 \cdot \text{Fluency} + 0.1 \cdot \text{Conciseness} ]
4) If you want a single number
If you need one output score, define it based on your priorities:
- News summarization: faithfulness and coverage matter most
- Meeting notes: coverage and structure matter more
- Creative summarization: fluency and style may matter more
If you want, I can help you design:
- a numeric scoring rubric
- a Python evaluation pipeline
- or a custom metric for your specific summarization task.