Prompt

llm output quality scoring rubric

Artificial Intelligence / AI Developer Tools1 observationLast seen Oct 9, 2026

Latest observation

Oct 9, 2026OpenAI APIWeb search: off

Here’s a practical LLM output quality scoring rubric you can use for evals, human review, or automated judging.

1) Core dimensions

Score each dimension on a 1–5 scale:

A. Relevance / Task Fit

How well does the output address the user’s request?

  • 5: Directly answers the question/task, no drift
  • 4: Mostly on target, minor omissions
  • 3: Partially relevant, some drift or incompleteness
  • 2: Weakly related or misses key intent
  • 1: Off-topic

B. Correctness / Factual Accuracy

How accurate are the facts, logic, code, or reasoning?

  • 5: Fully correct, no evident errors
  • 4: Mostly correct, minor issues that don’t matter much
  • 3: Mixed accuracy; notable errors or unsupported claims
  • 2: Major errors
  • 1: Mostly incorrect or hallucinated

C. Completeness

Does it include the needed parts?

  • 5: Fully complete
  • 4: Mostly complete; small missing detail
  • 3: Important gaps
  • 2: Many missing elements
  • 1: Severely incomplete

D. Clarity / Readability

Is it easy to understand?

  • 5: Clear, well-structured, concise
  • 4: Mostly clear; minor awkwardness
  • 3: Understandable but somewhat unclear
  • 2: Hard to follow
  • 1: Unclear or incoherent

E. Instruction Following

Does it obey formatting, style, constraints, and user preferences?

  • 5: Fully complies
  • 4: One small deviation
  • 3: Several deviations or partial compliance
  • 2: Major noncompliance
  • 1: Ignores instructions

F. Safety / Policy Compliance

Does it avoid unsafe, harmful, disallowed, or privacy-violating content?

  • 5: Fully safe/compliant
  • 3: Mild safety concerns or borderline content
  • 1: Clearly unsafe or disallowed

G. Helpfulness / Utility

Does the answer actually help accomplish the goal?

  • 5: Highly actionable and useful
  • 4: Useful with small limitations
  • 3: Some value, but not enough to fully solve task
  • 2: Low utility
  • 1: Not useful

2) Example weighted scoring

A simple weighted formula:

Overall Score =

  • Relevance: 20%
  • Correctness: 25%
  • Completeness: 20%
  • Clarity: 15%
  • Instruction Following: 10%
  • Safety: 5%
  • Helpfulness: 5%

Convert each 1–5 score to a 0–100 scale:

  • 1 = 0
  • 2 = 25
  • 3 = 50
  • 4 = 75
  • 5 = 100

Then compute weighted average.


3) Simple pass/fail gate

Before averaging, apply hard checks:

  • Fail if:
    • major factual errors in high-stakes domains
    • unsafe/disallowed content
    • ignores critical user constraints
    • produces irrelevant answer

This is useful because a high average score can hide a dangerous flaw.


4) Suggested rubric labels

You can map overall scores to labels:

  • 90–100: Excellent
  • 75–89: Good
  • 60–74: Fair
  • 40–59: Poor
  • 0–39: Fail

5) Short evaluator checklist

Use these yes/no prompts:

  • Did it answer the actual question?
  • Is it factually correct?
  • Did it include all required elements?
  • Is it clear and easy to read?
  • Did it follow format/style constraints?
  • Is it safe and policy-compliant?
  • Is it useful in practice?

6) Template you can copy

LLM Output Quality Rubric

1. Relevance (1-5): __
2. Correctness (1-5): __
3. Completeness (1-5): __
4. Clarity (1-5): __
5. Instruction Following (1-5): __
6. Safety/Compliance (1-5): __
7. Helpfulness (1-5): __

Overall Score (0-100): __
Label: Excellent / Good / Fair / Poor / Fail

Notes:
- Strengths:
- Weaknesses:
- Critical errors (if any):

If you want, I can also provide:

  1. a more rigorous rubric for human annotation,
  2. an automated eval rubric for LLM-as-judge, or
  3. a specialized rubric for code, summarization, or customer support.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.