Prompt

annotation guidelines inter-rater reliability llm

Artificial Intelligence / AI Developer Tools1 observationLast seen Oct 9, 2026

Latest observation

Oct 9, 2026OpenAI APIWeb search: off

For annotation guidelines in an inter-rater reliability (IRR) study with LLM-based annotation, the goal is to make labels as consistent, objective, and reproducible as possible.

Core principles

  1. Define the task precisely

    • State exactly what is being annotated.
    • Specify the unit of annotation: sentence, span, document, dialogue turn, etc.
    • Define the label set with mutually exclusive options if possible.
  2. Operationalize each label

    • For every class, give:
      • a short definition
      • inclusion criteria
      • exclusion criteria
      • edge-case handling
      • 2–5 examples
    • Avoid vague labels like “good,” “bad,” or “relevant” unless clearly defined.
  3. Make decisions independent of model output

    • If humans are annotating LLM outputs, guidelines should say whether to judge:
      • correctness,
      • factuality,
      • harm,
      • style,
      • helpfulness,
      • alignment with instruction,
      • etc.
    • Do not mix multiple constructs in one label unless you intend to measure a composite.
  4. Handle ambiguity explicitly

    • Add rules for:
      • insufficient context
      • multiple plausible labels
      • partial matches
      • overlapping spans
      • implicit vs explicit mentions
    • Decide whether annotators should use:
      • “uncertain”
      • “other”
      • “cannot determine”
  5. Prioritize consistency over perfection

    • If a case is borderline, choose a deterministic rule.
    • State tie-breakers:
      • “If both labels apply, choose X.”
      • “If the issue is present but not central, choose Y.”
    • This improves IRR a lot.

Recommended structure for guidelines

A good annotation guide usually includes:

  • Purpose of the annotation
  • Annotation unit
  • Label definitions
  • Decision rules
  • Examples and counterexamples
  • Edge cases
  • Quality control instructions
  • Escalation policy for uncertain cases

For LLM-specific IRR

If annotating LLM outputs, consider adding:

  • Ground truth source
    • Human reference? Retrieval source? External knowledge?
  • Tolerance for paraphrase
    • Is semantic equivalence enough?
  • Hallucination policy
    • What counts as unsupported?
  • Safety policy
    • What counts as harmful, unsafe, or policy-violating?
  • Prompt sensitivity
    • Should annotators ignore prompt quirks and judge final output only?

To improve inter-rater reliability

  • Run a pilot annotation round
  • Compute IRR on a small batch
  • Discuss disagreements
  • Revise guidelines
  • Repeat until agreement stabilizes

Useful metrics:

  • Cohen’s kappa for 2 annotators
  • Fleiss’ kappa for >2 annotators
  • Krippendorff’s alpha for flexible setups and missing data
  • Percent agreement as a simple but limited measure

Example guideline snippet

Label: Supported

  • The output is fully supported by the provided source/context.
  • Minor paraphrases are allowed.
  • Do not label as supported if any key claim is unsupported.

Label: Unsupported

  • At least one substantive claim is not supported by the source/context.
  • Minor stylistic additions do not count.

Label: Unclear

  • Use when the source is insufficient to determine support.

If you want, I can also give you:

  1. a template annotation guideline document,
  2. a rubric for LLM output evaluation, or
  3. best practices for improving IRR with examples.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.