Prompt
How can I score safety, factuality, and relevance without manually reading everything?
Latest observation
Oct 9, 2026
OpenAI APIWeb search: off
You can automate most of it with a triage pipeline that gives each response three scores: safety, factuality, and relevance. The key is to combine rules, model-based judges, and sampling, so humans only inspect the hardest cases.
1) Safety score
Use a classifier or policy model to flag:
- self-harm, violence, hate, harassment
- sexual content, especially minors
- illegal wrongdoing
- privacy / personal data leakage
- medical, legal, financial harmful advice
- prompt injection / jailbreak behavior
Practical scoring
Assign:
- 0 = clearly safe
- 1 = borderline / needs review
- 2 = unsafe
You can implement this with:
- a moderation API
- a fine-tuned safety classifier
- an LLM judge with a safety rubric
2) Factuality score
For factuality, don’t rely on surface similarity alone. Use claim checking:
- Split the answer into atomic claims.
- For each claim, retrieve supporting evidence from trusted sources.
- Score each claim as:
- supported
- refuted
- not enough evidence
Then aggregate:
- factuality = supported claims / total claims
Helpful shortcuts
- For RAG systems, compare answer claims against retrieved context.
- Use citation checking: “Does the cited source actually support this statement?”
- Penalize unsupported specifics, dates, numbers, and named entities.
3) Relevance score
Measure whether the answer actually addresses the prompt.
Possible signals:
- semantic similarity between prompt and answer
- coverage of requested subparts
- whether it follows instructions and format
- whether it avoids off-topic filler
Simple rubric
- 0 = off-topic
- 1 = partially answers
- 2 = directly and fully answers
4) Use an LLM as a judge
A strong approach is to ask another model to grade outputs with a rubric.
Example prompt to judge:
- Rate safety, factuality, and relevance from 1–5
- Require brief evidence for each score
- Force the judge to cite the exact text span causing the rating
This works well if you:
- keep the rubric strict and specific
- use multiple judges or multiple passes
- calibrate against a small human-labeled set
5) Reduce manual review with confidence thresholds
Only send to humans when:
- safety score is borderline
- factuality is low or uncertain
- relevance is low
- judges disagree
For example:
- auto-accept if all scores are high
- auto-reject if safety is clearly bad
- human review for middle cases
6) Best practice: combine all three
A solid production setup is:
- Safety classifier filters harmful content
- Retriever + claim checker evaluates factuality
- Relevance judge checks task completion
- Uncertainty routing sends only ambiguous cases to humans
7) If you want a quick implementation
A lightweight version:
- Use a moderation API for safety
- Use an LLM judge with a structured rubric for factuality and relevance
- Calibrate on 100–500 manually reviewed examples
- Track precision/recall, not just average scores
Example output format
{
"safety": 5,
"factuality": 3,
"relevance": 4,
"flags": ["unsupported medical claim", "minor off-topic detail"],
"confidence": 0.82
}
If you want, I can also give you:
- a rubric template,
- a judge prompt, or
- a Python evaluation pipeline for this.