Prompt
Can you recommend how to set up annotation guidelines so different reviewers score LLM outputs consistently?
Latest observation
Yes — the key is to make the rubric behaviorally specific, decision-oriented, and hard to interpret differently. Here’s a practical way to set up annotation guidelines so reviewers score LLM outputs more consistently.
1) Start with a small number of clearly defined dimensions
Avoid overly broad labels like “good” or “helpful.” Break evaluation into 3–6 dimensions max, such as:
- Correctness / factuality
- Completeness
- Relevance
- Instruction following
- Clarity / readability
- Safety / policy compliance
If you need a single overall score, compute it from these dimensions rather than asking reviewers to intuitively judge “quality.”
2) Write operational definitions for each dimension
For each label, define exactly what it means and what it does not mean.
Example:
Correctness
- The response is factually accurate based on the provided context or widely accepted facts.
- Minor stylistic issues do not count against this score.
- If the answer includes one major factual error, it cannot receive the top score.
Avoid vague terms like “mostly right” unless you define them with examples.
3) Use anchored rating scales
A 1–5 scale is more reliable if each point has a concrete anchor.
Example for correctness:
- 5 = No factual errors; fully supported.
- 4 = One minor error or imprecision that does not affect the main answer.
- 3 = Mixed: some correct content, but at least one meaningful error or omission.
- 2 = Mostly incorrect or unsupported.
- 1 = Completely incorrect or hallucinated.
Anchors reduce reviewer drift and make disagreements easier to diagnose.
4) Add decision rules and precedence
Reviewers often disagree because they are making different tradeoffs. Spell those out.
Examples:
- If the response is fluent but factually wrong, score correctness based on truthfulness, not writing quality.
- If the response refuses appropriately, do not penalize it for lacking a direct answer.
- If the response violates safety policy, safety score overrides helpfulness.
Use precedence rules like:
- Safety
- Correctness
- Instruction following
- Style
5) Include positive and negative examples for each score
Examples are one of the best ways to align reviewers.
For each dimension, show:
- A strong example
- A borderline example
- A failing example
Make sure examples cover common ambiguities, not just obvious cases.
6) Define how to handle edge cases
Most inconsistency comes from edge cases. Explicitly address them:
- Partial answers
- Refusals
- Mixed correctness
- Overly verbose but correct responses
- Answers that are correct but don’t follow format
- Unsupported claims with hedging
- Ambiguous prompts
- Multiple valid answers
A short “if X, then Y” section helps a lot.
7) Separate “objective” and “subjective” judgments
If possible, distinguish between:
- Objective checks: factual errors, policy violations, format compliance
- Subjective judgments: clarity, conciseness, tone
This reduces arguments because reviewers know which parts require strict adherence and which allow judgment.
8) Make the unit of evaluation explicit
Tell reviewers exactly what they are judging:
- Just the final answer?
- The whole conversation?
- The answer relative to retrieved context?
- The answer relative to a source document?
Without this, people will score the same output differently depending on whether they mentally include prior turns or external knowledge.
9) Calibrate with a shared training set
Before real annotation:
- Have all reviewers score the same 20–50 examples.
- Compare results.
- Discuss disagreements.
- Update the guidelines.
- Repeat until agreement improves.
This calibration step often matters more than the written rubric itself.
10) Measure agreement and inspect disagreement patterns
Use inter-annotator agreement metrics appropriate to your setup:
- Cohen’s kappa for 2 annotators
- Fleiss’ kappa or Krippendorff’s alpha for more
- Also track raw percent agreement
Then inspect where disagreements occur:
- One reviewer lenient on omissions?
- One reviewer harsher on tone?
- One reviewer treating uncertainty differently?
Update guidelines based on those patterns.
11) Use a “gold” set and periodic QA
Keep a small set of expert-labeled items as references.
- Use them in onboarding
- Insert them periodically into live annotation
- Flag annotators who drift
This helps maintain consistency over time.
12) Keep the rubric short enough to use
A perfect rubric that nobody follows is worse than a simpler one people can apply consistently. If reviewers need to make too many nuanced judgments, simplify:
- Fewer categories
- Fewer score levels
- More examples
- More hard rules
A simple template you can use
Dimension: Correctness
Definition: The response must be factually accurate and supported by the prompt/context.
Score 5: Fully correct; no material errors.
Score 4: Minor imprecision; core answer correct.
Score 3: Some correct content, but at least one material error.
Score 2: Mostly incorrect or unsupported.
Score 1: Completely incorrect.
Rules:
- Ignore grammar unless it changes meaning.
- If the model refuses due to safety, do not score correctness unless it makes factual claims.
- If there are multiple claims, score based on the most serious claim error.
Examples:
Include 2–3 per score level.
Practical tips that help a lot
- Prefer binary or 3-point judgments when possible
- Require reviewers to cite the text span causing the score
- Make disagreements resolvable by evidence, not taste
- Version-control the rubric and examples
- Update the guidelines only in batches, not ad hoc
If you want, I can also draft:
- a full annotation guideline template,
- a rubric for a specific use case like helpfulness/factuality/safety, or
- a reviewer training workflow for improving inter-annotator agreement.