Prompt
llm output quality scoring rubric
Latest observation
Here’s a practical LLM output quality scoring rubric you can use for evals, human review, or automated judging.
1) Core dimensions
Score each dimension on a 1–5 scale:
A. Relevance / Task Fit
How well does the output address the user’s request?
- 5: Directly answers the question/task, no drift
- 4: Mostly on target, minor omissions
- 3: Partially relevant, some drift or incompleteness
- 2: Weakly related or misses key intent
- 1: Off-topic
B. Correctness / Factual Accuracy
How accurate are the facts, logic, code, or reasoning?
- 5: Fully correct, no evident errors
- 4: Mostly correct, minor issues that don’t matter much
- 3: Mixed accuracy; notable errors or unsupported claims
- 2: Major errors
- 1: Mostly incorrect or hallucinated
C. Completeness
Does it include the needed parts?
- 5: Fully complete
- 4: Mostly complete; small missing detail
- 3: Important gaps
- 2: Many missing elements
- 1: Severely incomplete
D. Clarity / Readability
Is it easy to understand?
- 5: Clear, well-structured, concise
- 4: Mostly clear; minor awkwardness
- 3: Understandable but somewhat unclear
- 2: Hard to follow
- 1: Unclear or incoherent
E. Instruction Following
Does it obey formatting, style, constraints, and user preferences?
- 5: Fully complies
- 4: One small deviation
- 3: Several deviations or partial compliance
- 2: Major noncompliance
- 1: Ignores instructions
F. Safety / Policy Compliance
Does it avoid unsafe, harmful, disallowed, or privacy-violating content?
- 5: Fully safe/compliant
- 3: Mild safety concerns or borderline content
- 1: Clearly unsafe or disallowed
G. Helpfulness / Utility
Does the answer actually help accomplish the goal?
- 5: Highly actionable and useful
- 4: Useful with small limitations
- 3: Some value, but not enough to fully solve task
- 2: Low utility
- 1: Not useful
2) Example weighted scoring
A simple weighted formula:
Overall Score =
- Relevance: 20%
- Correctness: 25%
- Completeness: 20%
- Clarity: 15%
- Instruction Following: 10%
- Safety: 5%
- Helpfulness: 5%
Convert each 1–5 score to a 0–100 scale:
- 1 = 0
- 2 = 25
- 3 = 50
- 4 = 75
- 5 = 100
Then compute weighted average.
3) Simple pass/fail gate
Before averaging, apply hard checks:
- Fail if:
- major factual errors in high-stakes domains
- unsafe/disallowed content
- ignores critical user constraints
- produces irrelevant answer
This is useful because a high average score can hide a dangerous flaw.
4) Suggested rubric labels
You can map overall scores to labels:
- 90–100: Excellent
- 75–89: Good
- 60–74: Fair
- 40–59: Poor
- 0–39: Fail
5) Short evaluator checklist
Use these yes/no prompts:
- Did it answer the actual question?
- Is it factually correct?
- Did it include all required elements?
- Is it clear and easy to read?
- Did it follow format/style constraints?
- Is it safe and policy-compliant?
- Is it useful in practice?
6) Template you can copy
LLM Output Quality Rubric
1. Relevance (1-5): __
2. Correctness (1-5): __
3. Completeness (1-5): __
4. Clarity (1-5): __
5. Instruction Following (1-5): __
6. Safety/Compliance (1-5): __
7. Helpfulness (1-5): __
Overall Score (0-100): __
Label: Excellent / Good / Fair / Poor / Fail
Notes:
- Strengths:
- Weaknesses:
- Critical errors (if any):
If you want, I can also provide:
- a more rigorous rubric for human annotation,
- an automated eval rubric for LLM-as-judge, or
- a specialized rubric for code, summarization, or customer support.