Prompt
annotation guidelines inter-rater reliability llm
Latest observation
Oct 9, 2026
OpenAI APIWeb search: off
For annotation guidelines in an inter-rater reliability (IRR) study with LLM-based annotation, the goal is to make labels as consistent, objective, and reproducible as possible.
Core principles
-
Define the task precisely
- State exactly what is being annotated.
- Specify the unit of annotation: sentence, span, document, dialogue turn, etc.
- Define the label set with mutually exclusive options if possible.
-
Operationalize each label
- For every class, give:
- a short definition
- inclusion criteria
- exclusion criteria
- edge-case handling
- 2–5 examples
- Avoid vague labels like “good,” “bad,” or “relevant” unless clearly defined.
- For every class, give:
-
Make decisions independent of model output
- If humans are annotating LLM outputs, guidelines should say whether to judge:
- correctness,
- factuality,
- harm,
- style,
- helpfulness,
- alignment with instruction,
- etc.
- Do not mix multiple constructs in one label unless you intend to measure a composite.
- If humans are annotating LLM outputs, guidelines should say whether to judge:
-
Handle ambiguity explicitly
- Add rules for:
- insufficient context
- multiple plausible labels
- partial matches
- overlapping spans
- implicit vs explicit mentions
- Decide whether annotators should use:
- “uncertain”
- “other”
- “cannot determine”
- Add rules for:
-
Prioritize consistency over perfection
- If a case is borderline, choose a deterministic rule.
- State tie-breakers:
- “If both labels apply, choose X.”
- “If the issue is present but not central, choose Y.”
- This improves IRR a lot.
Recommended structure for guidelines
A good annotation guide usually includes:
- Purpose of the annotation
- Annotation unit
- Label definitions
- Decision rules
- Examples and counterexamples
- Edge cases
- Quality control instructions
- Escalation policy for uncertain cases
For LLM-specific IRR
If annotating LLM outputs, consider adding:
- Ground truth source
- Human reference? Retrieval source? External knowledge?
- Tolerance for paraphrase
- Is semantic equivalence enough?
- Hallucination policy
- What counts as unsupported?
- Safety policy
- What counts as harmful, unsafe, or policy-violating?
- Prompt sensitivity
- Should annotators ignore prompt quirks and judge final output only?
To improve inter-rater reliability
- Run a pilot annotation round
- Compute IRR on a small batch
- Discuss disagreements
- Revise guidelines
- Repeat until agreement stabilizes
Useful metrics:
- Cohen’s kappa for 2 annotators
- Fleiss’ kappa for >2 annotators
- Krippendorff’s alpha for flexible setups and missing data
- Percent agreement as a simple but limited measure
Example guideline snippet
Label: Supported
- The output is fully supported by the provided source/context.
- Minor paraphrases are allowed.
- Do not label as supported if any key claim is unsupported.
Label: Unsupported
- At least one substantive claim is not supported by the source/context.
- Minor stylistic additions do not count.
Label: Unclear
- Use when the source is insufficient to determine support.
If you want, I can also give you:
- a template annotation guideline document,
- a rubric for LLM output evaluation, or
- best practices for improving IRR with examples.