Prompt
I need evaluation for summarization that handles partial correctness
Latest observation
If you want to evaluate summarization with partial correctness, the best approach is to use a metric or rubric that gives graded credit instead of only all-or-nothing scoring.
Good evaluation options
1. ROUGE with semantic/manual adjustment
ROUGE is common, but it mostly checks overlap with a reference summary.
It can give partial credit when the summary captures some key words or phrases, but it does not fully capture meaning.
Use:
- ROUGE-1 for unigram overlap
- ROUGE-2 for phrase overlap
- ROUGE-L for sequence similarity
Best when:
- You have reference summaries
- You want a quick automatic baseline
Limitation:
- Misses paraphrases and deeper meaning
2. Fact-based or content-unit scoring
Break the source/reference into atomic facts or content units, then score the generated summary by how many units are:
- correct
- partially correct
- incorrect
- missing
Example scoring:
- correct = 1.0
- partially correct = 0.5
- incorrect = 0
- hallucinated = -1 or 0 depending on setup
This is good when you want explicit partial credit.
3. Pyramid method / SCU-based evaluation
The Pyramid Method evaluates summaries by comparing against Summary Content Units (SCUs).
A summary gets partial credit for each SCU it includes, even if wording differs.
This is one of the most established methods for partial correctness.
Best when:
- You need a human evaluation
- You care about content coverage, not just wording
4. Semantic similarity metrics
These compare meaning rather than exact overlap:
- BERTScore
- BLEURT
- MoverScore
- BARTScore
These are better for paraphrases and partial matches.
Use them when:
- You want automated semantic evaluation
- You need partial credit for meaning-preserving rewrites
5. LLM-as-judge with rubric-based scoring
You can ask an LLM to score summaries using a rubric like:
- 0 = incorrect
- 1 = partially correct
- 2 = mostly correct
- 3 = fully correct
And separately score:
- factual accuracy
- coverage
- coherence
- conciseness
This is flexible and often works well for partial correctness if the rubric is clear.
Recommended setup
If your goal is specifically partial correctness, I’d suggest:
-
Define a rubric
- exact match
- partially correct
- incorrect
- hallucinated
-
Use content units or atomic facts
- score each fact individually
-
Combine with semantic metrics
- BERTScore or ROUGE as a baseline
-
Report both
- fact coverage
- factual precision
- partial-credit score
Example partial-credit scoring
Suppose a reference summary has 5 key facts:
- 3 facts fully correct = 3 points
- 1 fact partially correct = 0.5 points
- 1 fact missing = 0 points
Total = 3.5 / 5 = 70%
You can also define:
- Precision: how many stated facts are correct
- Recall: how many reference facts are covered
- F1: balance of both
Partial correctness can be built into the scoring of each fact.
If you want a single practical metric
Use:
- BERTScore for automatic partial semantic similarity
- manual fact-level rubric for high-quality evaluation
If you want, I can also give you:
- a formal evaluation rubric,
- a Python implementation, or
- a table comparing metrics for partial correctness.