Prompt

How do I score summarization quality with human and automated evals?

Artificial Intelligence / AI Developer Tools1 observationLast seen Oct 10, 2026

Latest observation

Oct 10, 2026OpenAI APIWeb search: off

Scoring summarization quality usually works best if you combine human evaluation and automated evaluation, because they measure different things:

  • Humans catch faithfulness, usefulness, and subtle errors.
  • Automated metrics give fast, cheap, repeatable signals at scale.

A good setup is to score along a few dimensions rather than one single “summary quality” number.

1) What to evaluate

Common summarization dimensions:

  1. Faithfulness / factuality

    • Is every claim supported by the source?
    • Are there hallucinations, distortions, or omissions of key facts?
  2. Coverage / content selection

    • Does the summary include the main points?
    • Does it miss important information?
  3. Conciseness

    • Is it short and not repetitive?
    • Does it avoid unnecessary detail?
  4. Coherence / readability

    • Is it well-structured and easy to understand?
    • Are pronouns, entities, and transitions clear?
  5. Relevance / focus

    • Is the summary focused on the source topic and user need?

If you only have room for a few dimensions, prioritize faithfulness and coverage.


2) Human evaluation: how to score it

A. Likert-scale rating

Have annotators rate each summary on a scale, e.g. 1–5:

  • Faithfulness

    • 1 = many factual errors
    • 3 = minor issues
    • 5 = fully supported by source
  • Coverage

    • 1 = misses most key points
    • 3 = covers some important points
    • 5 = covers nearly all important points
  • Coherence

    • 1 = hard to read
    • 5 = clear and well organized

This is simple and works well if you want aggregate scores.

B. Pairwise preference

Show annotators two summaries for the same source and ask which is better, or whether they are tied.

This often yields more reliable judgments than absolute ratings.

Use it for:

  • model comparisons
  • product A/B tests
  • ranking different prompt/model versions

C. Error annotation

Instead of just rating, ask annotators to mark errors:

  • hallucinated fact
  • missing key fact
  • wrong entity/date/number
  • redundant content
  • unclear reference

This helps diagnose failures and improve the system.


3) Human eval best practices

  • Use a clear rubric. Define each score level with examples.
  • Train annotators. Do a calibration round with sample summaries.
  • Measure agreement. Use Krippendorff’s alpha, Cohen’s kappa, or pairwise agreement.
  • Blind annotators. Hide the system name/model version.
  • Randomize order. Avoid position bias in pairwise comparisons.
  • Sample representative data. Include easy, hard, and edge-case examples.
  • Annotate source access. For faithfulness, annotators should compare summary to source text.

A practical workflow:

  1. Randomly sample documents.
  2. Generate summaries from all systems.
  3. Have 2–3 annotators rate each summary.
  4. Average scores or use majority vote.
  5. Report confidence intervals.

4) Automated evaluation: common metrics

A. Lexical overlap metrics

These compare summary words to reference summaries.

  • ROUGE-1 / ROUGE-2 / ROUGE-L
    • Measures unigram, bigram, and sequence overlap.
    • Widely used and easy to compute.

Pros:

  • simple
  • cheap
  • standard baseline

Cons:

  • misses paraphrases
  • doesn’t measure factuality well
  • can reward copying rather than good abstraction

Best for:

  • comparison against reference summaries
  • quick baseline checks

B. Semantic similarity metrics

These use embeddings or learned models to compare meaning.

  • BERTScore
  • MoverScore
  • other embedding-based metrics

Pros:

  • better for paraphrase matching
  • more semantic than ROUGE

Cons:

  • still not a direct factuality measure
  • may correlate imperfectly with human judgment

C. Factuality / consistency metrics

These try to detect whether the summary is supported by the source.

Examples:

  • SummaC
  • FactCC
  • QAFactEval
  • BARTScore in some setups
  • question-answering based factuality checks

Pros:

  • more aligned with hallucination detection
  • useful for source-grounded summarization

Cons:

  • can be brittle
  • often domain-sensitive
  • may miss subtle factual errors

D. LLM-based evals

Use a strong language model as a judge with a rubric.

Pros:

  • flexible
  • can evaluate multiple dimensions
  • can give explanations

Cons:

  • can be biased
  • needs careful prompting and calibration
  • should be validated against humans

This is often useful when you need scalable qualitative scoring.


5) A practical scoring framework

A strong setup is:

Human scores

Rate each summary on 1–5 for:

  • Faithfulness
  • Coverage
  • Coherence
  • Conciseness

Then compute:

  • mean score per dimension
  • overall weighted score, e.g.:
    • 40% faithfulness
    • 30% coverage
    • 20% coherence
    • 10% conciseness

Automated scores

Track:

  • ROUGE for overlap
  • BERTScore for semantic similarity
  • a factuality metric for hallucination detection

Then correlate these with human scores to see which metrics are meaningful for your task.


6) Example scoring rubric

Here’s a simple faithfulness rubric:

  • 5: All statements are supported by the source; no hallucinations.
  • 4: Mostly faithful; one minor unsupported detail.
  • 3: Some unsupported or distorted content, but main message is correct.
  • 2: Several factual errors or misleading statements.
  • 1: Mostly hallucinated or contradicted by source.

Coverage rubric:

  • 5: Includes all major points.
  • 4: Misses one minor point.
  • 3: Captures some major points but misses others.
  • 2: Captures only a small subset.
  • 1: Misses almost all important information.

7) How to combine human and automated evals

A good pattern is:

  1. Use automated metrics for iteration

    • During development, use ROUGE/BERTScore/factuality checks to quickly compare versions.
  2. Use human eval for validation

    • Before release, confirm quality with human judgments.
  3. Calibrate automated metrics

    • Check correlation with human ratings.
    • If a metric tracks human faithfulness poorly, don’t rely on it for gating.
  4. Use human eval on a held-out benchmark

    • Keep a fixed evaluation set to compare future model versions.

8) What not to do

  • Don’t rely on ROUGE alone.
  • Don’t average everything into one score without checking each dimension.
  • Don’t use LLM judges without validating them against humans.
  • Don’t evaluate only on easy examples or only on one domain.
  • Don’t ignore inter-annotator agreement.

9) Recommended minimal setup

If you want a simple, robust process:

  • Human eval

    • 2 annotators
    • 1–5 ratings for faithfulness, coverage, coherence
    • pairwise preference for model comparisons
  • Automated eval

    • ROUGE
    • BERTScore
    • a factuality metric or LLM judge
  • Report

    • mean scores
    • agreement
    • examples of failure cases
    • correlation between automatic and human scores

If you want, I can also give you:

  1. a ready-to-use annotation rubric,
  2. a Python evaluation script, or
  3. a template for an LLM-as-judge prompt.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.