Prompt

Can you recommend how to set up annotation guidelines so different reviewers score LLM outputs consistently?

Artificial Intelligence / AI Developer Tools1 observationLast seen Oct 10, 2026

Latest observation

Oct 10, 2026OpenAI APIWeb search: off

Yes — the key is to make the rubric behaviorally specific, decision-oriented, and hard to interpret differently. Here’s a practical way to set up annotation guidelines so reviewers score LLM outputs more consistently.

1) Start with a small number of clearly defined dimensions

Avoid overly broad labels like “good” or “helpful.” Break evaluation into 3–6 dimensions max, such as:

  • Correctness / factuality
  • Completeness
  • Relevance
  • Instruction following
  • Clarity / readability
  • Safety / policy compliance

If you need a single overall score, compute it from these dimensions rather than asking reviewers to intuitively judge “quality.”

2) Write operational definitions for each dimension

For each label, define exactly what it means and what it does not mean.

Example:

Correctness

  • The response is factually accurate based on the provided context or widely accepted facts.
  • Minor stylistic issues do not count against this score.
  • If the answer includes one major factual error, it cannot receive the top score.

Avoid vague terms like “mostly right” unless you define them with examples.

3) Use anchored rating scales

A 1–5 scale is more reliable if each point has a concrete anchor.

Example for correctness:

  • 5 = No factual errors; fully supported.
  • 4 = One minor error or imprecision that does not affect the main answer.
  • 3 = Mixed: some correct content, but at least one meaningful error or omission.
  • 2 = Mostly incorrect or unsupported.
  • 1 = Completely incorrect or hallucinated.

Anchors reduce reviewer drift and make disagreements easier to diagnose.

4) Add decision rules and precedence

Reviewers often disagree because they are making different tradeoffs. Spell those out.

Examples:

  • If the response is fluent but factually wrong, score correctness based on truthfulness, not writing quality.
  • If the response refuses appropriately, do not penalize it for lacking a direct answer.
  • If the response violates safety policy, safety score overrides helpfulness.

Use precedence rules like:

  1. Safety
  2. Correctness
  3. Instruction following
  4. Style

5) Include positive and negative examples for each score

Examples are one of the best ways to align reviewers.

For each dimension, show:

  • A strong example
  • A borderline example
  • A failing example

Make sure examples cover common ambiguities, not just obvious cases.

6) Define how to handle edge cases

Most inconsistency comes from edge cases. Explicitly address them:

  • Partial answers
  • Refusals
  • Mixed correctness
  • Overly verbose but correct responses
  • Answers that are correct but don’t follow format
  • Unsupported claims with hedging
  • Ambiguous prompts
  • Multiple valid answers

A short “if X, then Y” section helps a lot.

7) Separate “objective” and “subjective” judgments

If possible, distinguish between:

  • Objective checks: factual errors, policy violations, format compliance
  • Subjective judgments: clarity, conciseness, tone

This reduces arguments because reviewers know which parts require strict adherence and which allow judgment.

8) Make the unit of evaluation explicit

Tell reviewers exactly what they are judging:

  • Just the final answer?
  • The whole conversation?
  • The answer relative to retrieved context?
  • The answer relative to a source document?

Without this, people will score the same output differently depending on whether they mentally include prior turns or external knowledge.

9) Calibrate with a shared training set

Before real annotation:

  1. Have all reviewers score the same 20–50 examples.
  2. Compare results.
  3. Discuss disagreements.
  4. Update the guidelines.
  5. Repeat until agreement improves.

This calibration step often matters more than the written rubric itself.

10) Measure agreement and inspect disagreement patterns

Use inter-annotator agreement metrics appropriate to your setup:

  • Cohen’s kappa for 2 annotators
  • Fleiss’ kappa or Krippendorff’s alpha for more
  • Also track raw percent agreement

Then inspect where disagreements occur:

  • One reviewer lenient on omissions?
  • One reviewer harsher on tone?
  • One reviewer treating uncertainty differently?

Update guidelines based on those patterns.

11) Use a “gold” set and periodic QA

Keep a small set of expert-labeled items as references.

  • Use them in onboarding
  • Insert them periodically into live annotation
  • Flag annotators who drift

This helps maintain consistency over time.

12) Keep the rubric short enough to use

A perfect rubric that nobody follows is worse than a simpler one people can apply consistently. If reviewers need to make too many nuanced judgments, simplify:

  • Fewer categories
  • Fewer score levels
  • More examples
  • More hard rules

A simple template you can use

Dimension: Correctness

Definition: The response must be factually accurate and supported by the prompt/context.
Score 5: Fully correct; no material errors.
Score 4: Minor imprecision; core answer correct.
Score 3: Some correct content, but at least one material error.
Score 2: Mostly incorrect or unsupported.
Score 1: Completely incorrect.

Rules:

  • Ignore grammar unless it changes meaning.
  • If the model refuses due to safety, do not score correctness unless it makes factual claims.
  • If there are multiple claims, score based on the most serious claim error.

Examples:
Include 2–3 per score level.

Practical tips that help a lot

  • Prefer binary or 3-point judgments when possible
  • Require reviewers to cite the text span causing the score
  • Make disagreements resolvable by evidence, not taste
  • Version-control the rubric and examples
  • Update the guidelines only in batches, not ad hoc

If you want, I can also draft:

  1. a full annotation guideline template,
  2. a rubric for a specific use case like helpfulness/factuality/safety, or
  3. a reviewer training workflow for improving inter-annotator agreement.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.