Prompt

I need evaluation for summarization that handles partial correctness

Artificial Intelligence / AI Developer Tools1 observationLast seen Oct 9, 2026

Latest observation

Oct 9, 2026OpenAI APIWeb search: off

If you want to evaluate summarization with partial correctness, the best approach is to use a metric or rubric that gives graded credit instead of only all-or-nothing scoring.

Good evaluation options

1. ROUGE with semantic/manual adjustment

ROUGE is common, but it mostly checks overlap with a reference summary.
It can give partial credit when the summary captures some key words or phrases, but it does not fully capture meaning.

Use:

  • ROUGE-1 for unigram overlap
  • ROUGE-2 for phrase overlap
  • ROUGE-L for sequence similarity

Best when:

  • You have reference summaries
  • You want a quick automatic baseline

Limitation:

  • Misses paraphrases and deeper meaning

2. Fact-based or content-unit scoring

Break the source/reference into atomic facts or content units, then score the generated summary by how many units are:

  • correct
  • partially correct
  • incorrect
  • missing

Example scoring:

  • correct = 1.0
  • partially correct = 0.5
  • incorrect = 0
  • hallucinated = -1 or 0 depending on setup

This is good when you want explicit partial credit.


3. Pyramid method / SCU-based evaluation

The Pyramid Method evaluates summaries by comparing against Summary Content Units (SCUs).

A summary gets partial credit for each SCU it includes, even if wording differs.
This is one of the most established methods for partial correctness.

Best when:

  • You need a human evaluation
  • You care about content coverage, not just wording

4. Semantic similarity metrics

These compare meaning rather than exact overlap:

  • BERTScore
  • BLEURT
  • MoverScore
  • BARTScore

These are better for paraphrases and partial matches.

Use them when:

  • You want automated semantic evaluation
  • You need partial credit for meaning-preserving rewrites

5. LLM-as-judge with rubric-based scoring

You can ask an LLM to score summaries using a rubric like:

  • 0 = incorrect
  • 1 = partially correct
  • 2 = mostly correct
  • 3 = fully correct

And separately score:

  • factual accuracy
  • coverage
  • coherence
  • conciseness

This is flexible and often works well for partial correctness if the rubric is clear.


Recommended setup

If your goal is specifically partial correctness, I’d suggest:

  1. Define a rubric

    • exact match
    • partially correct
    • incorrect
    • hallucinated
  2. Use content units or atomic facts

    • score each fact individually
  3. Combine with semantic metrics

    • BERTScore or ROUGE as a baseline
  4. Report both

    • fact coverage
    • factual precision
    • partial-credit score

Example partial-credit scoring

Suppose a reference summary has 5 key facts:

  • 3 facts fully correct = 3 points
  • 1 fact partially correct = 0.5 points
  • 1 fact missing = 0 points

Total = 3.5 / 5 = 70%

You can also define:

  • Precision: how many stated facts are correct
  • Recall: how many reference facts are covered
  • F1: balance of both

Partial correctness can be built into the scoring of each fact.


If you want a single practical metric

Use:

  • BERTScore for automatic partial semantic similarity
  • manual fact-level rubric for high-quality evaluation

If you want, I can also give you:

  1. a formal evaluation rubric,
  2. a Python implementation, or
  3. a table comparing metrics for partial correctness.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.