Prompt
Can you help me design an eval plan for an LLM product that needs human review for edge cases but automated checks for most outputs?
Latest observation
Absolutely. A good eval plan for this kind of LLM product usually has three layers:
- Automated checks for the majority of outputs
- Human review for edge cases and ambiguous failures
- Ongoing monitoring in production to catch regressions and drift
Below is a practical template you can adapt.
1) Start with the product risk profile
Before defining metrics, decide what can go wrong.
Common risk categories
- Correctness: factual errors, wrong calculations, bad reasoning
- Policy/compliance: unsafe, disallowed, privacy leaks, legal issues
- Tone/style: rude, overly verbose, inconsistent with brand
- Completeness: missing required fields or steps
- Grounding: unsupported claims when the product should use provided context
- Tool use: bad function calls, malformed JSON, wrong API parameters
- User harm: advice causing financial, medical, or safety harm
Classify output severity
Use severity levels to decide what gets auto-checked vs human-reviewed:
- Sev 0: harmless style issues
- Sev 1: minor quality problems
- Sev 2: user-visible correctness issues
- Sev 3: high-risk issues, policy violations, harmful advice
- Sev 4: critical safety/legal/security issues
A useful rule:
- Automate everything that is stable, measurable, and low-risk
- Send to humans anything high-impact, ambiguous, or rare
2) Define evaluation dimensions
Pick a small set of dimensions that matter for your product.
Example dimensions
Core quality
- Task success
- Accuracy / factuality
- Instruction following
- Completeness
- Conciseness / verbosity control
Safety and compliance
- Policy adherence
- PII handling
- Harmful content avoidance
- Refusal quality for disallowed requests
Product-specific
- Schema validity for structured outputs
- Tool-call correctness
- Citation quality / grounding
- Brand voice
- Localization quality
For each dimension, define:
- What “good” means
- How it is measured
- Whether it is automated, human-reviewed, or both
- The failure threshold
3) Split evals into automated vs human review
Best candidates for automation
These are usually deterministic or machine-checkable:
- JSON / schema validity
- Required field presence
- Regex / formatting checks
- Tool-call argument validation
- Exact-match or fuzzy-match against known answers
- Citation presence and formatting
- Policy keyword or classifier checks
- PII detection
- Length limits
- Language detection
- Toxicity classifiers
- Unit tests for business rules
Best candidates for human review
These are subjective or high-stakes:
- Factual correctness in open-ended answers
- Reasoning quality
- Helpfulness / completeness
- Borderline safety cases
- Tone and empathy
- Appropriate refusal behavior
- Groundedness when evidence is nuanced
- Edge cases outside the training/eval set
4) Use a tiered review flow
A practical design is:
Tier A: Automated pass/fail
Run every output through checks like:
- schema validation
- safety filters
- policy classifiers
- known-answer scoring
- citation/grounding heuristics
- tool-call validation
If it fails a hard rule, mark as fail immediately.
Tier B: Uncertainty-based human review
Route outputs to humans if:
- automated checks conflict
- confidence is low
- output falls near decision thresholds
- prompt contains rare or risky intent
- the model self-reports uncertainty, if you use that signal
- the output is from a new prompt type or new model version
Tier C: Sampled human audit
Even for outputs that pass automation, sample a small percentage for human review to catch blind spots.
A common pattern:
- Review 100% of failures
- Review all high-risk categories
- Review 1–5% of routine successes
- Review all new use cases during launch
- Review all outputs from model changes during canary period
5) Build an eval dataset strategy
You need several dataset slices.
A. Golden set
A stable set of representative examples with expected outputs or rubric-based judgments. Use this for:
- regression testing
- model comparisons
- release gating
B. Edge-case set
Examples designed to break the system:
- adversarial prompts
- ambiguous instructions
- conflicting requirements
- malformed inputs
- rare user intents
- policy boundary cases
C. High-risk set
Focus on scenarios with biggest user or business impact:
- medical, legal, financial
- privacy-sensitive cases
- security-related requests
- regulated content
- tool-action requests with real consequences
D. Production sample set
A rolling sample of real traffic, anonymized if necessary. Use it to discover:
- unknown unknowns
- prompt distribution shift
- recurring failure patterns
6) Define scoring rubrics for humans
To make human review reliable, use a rubric rather than freeform judgment.
Example rubric
For each dimension:
- 0 = unacceptable
- 1 = poor
- 2 = acceptable
- 3 = good
- 4 = excellent
Or simpler:
- Pass / Fail / Needs escalation
Review instructions should include
- what to check
- what counts as a failure
- examples of borderline cases
- how to handle uncertainty
- when to escalate to a subject matter expert
Measure reviewer quality
To keep human evals trustworthy:
- use overlap on a subset of samples
- measure inter-rater agreement
- periodically calibrate reviewers
- include known-answer “gold” items in the review queue
7) Define release gates
Set explicit criteria for shipping.
Example release gates
A new model version ships only if:
- schema validity ≥ 99.5%
- disallowed content rate = 0 on high-risk set
- task success improves or stays within tolerance
- factual error rate does not regress more than X%
- no critical failures in human-reviewed edge cases
- no degradation in top business metrics
Use “must-pass” and “nice-to-have” buckets
- Must-pass: safety, compliance, schema correctness, tool reliability
- Nice-to-have: style, verbosity, minor helpfulness gains
This avoids over-optimizing subjective metrics at the expense of safety.
8) Include regression testing and versioning
Every prompt, model, retrieval index, tool schema, and policy update can change behavior.
Track versions of:
- model
- system prompt
- developer prompt
- tools
- retrieval corpus
- post-processing rules
- safety classifiers
For each release:
- run the golden and edge-case suites
- compare against baseline
- flag regressions by category
- keep a changelog of failures and fixes
9) Use targeted error analysis
After each eval cycle, group failures by type:
- hallucination
- missed constraint
- wrong format
- unsafe answer
- incorrect refusal
- poor citation
- tool misuse
Then identify:
- which are automatable
- which need new test cases
- which need prompt/model changes
- which need policy changes
- which need human escalation rules
This turns evals into a feedback loop instead of just a scorecard.
10) Add production monitoring
Even great offline evals miss real-world behavior.
Monitor:
- failure rates
- escalation rates
- user corrections / edits
- abandonment
- re-asks
- complaint rates
- safety-filter triggers
- tool error rates
- drift in prompt distribution
- performance by segment
Set alerts for spikes and regressions.
A useful pattern:
- automated dashboards for daily monitoring
- weekly human audit of sampled conversations
- monthly deep-dive on top failure modes
11) Suggested operating model
Here’s a simple, effective setup:
During development
- Run automated checks on every candidate
- Human review all failures and edge-case samples
- Iterate on prompts/rules/tests
During release
- Canary new versions on a small traffic slice
- Review all canary outputs that hit risk thresholds
- Compare against current version
In production
- Auto-check all outputs
- Route risky/uncertain outputs to human review
- Sample routine outputs for audit
- Feed failures back into the eval set
12) A concrete example architecture
You could implement something like:
-
Input classification
- detect topic, risk level, language, intent
-
Automated validators
- schema validator
- policy classifier
- grounding/citation check
- business-rule checks
-
Routing
- if hard fail → block or escalate
- if uncertain / high risk → human review
- else → release
-
Human review UI
- show prompt, output, retrieved evidence, tool traces
- collect rubric scores and failure tags
-
Dashboard
- aggregate metrics
- trend failures by version and category
- track reviewer agreement
13) Example metrics table
| Dimension | Method | Threshold | Owner |
|---|---|---|---|
| Schema validity | Automated | ≥ 99.5% | Eng |
| Policy violations | Automated + human audit | 0 critical | Safety |
| Task success | Human rubric | ≥ baseline | Product |
| Factuality | Human + spot checks | no regression | DS/PM |
| Tool-call validity | Automated | ≥ 99% | Eng |
| Escalation accuracy | Human review | improve over time | Ops |
| User satisfaction | Production metric | positive trend | Product |
14) Practical recommendation: decision matrix
Use this rule of thumb:
Auto-check if:
- the requirement is explicit
- the output format is structured
- a deterministic validator exists
- the cost of a false positive/negative is low
Human-review if:
- the output is open-ended
- correctness is subjective
- the category is rare
- mistakes are high-impact
- the model confidence is low
- the automated signals disagree
15) Minimal starting plan
If you want something lightweight to begin with:
- Create a golden set of 100–300 examples
- Add an edge-case set of 50–100 risky examples
- Build automated checks for format, policy, and known business rules
- Define a human rubric for accuracy/helpfulness/safety
- Review:
- 100% of automated failures
- 100% of high-risk examples
- 5% of routine passes
- Track failures by category and version
- Use release gates to block regressions
- Add production sampling and monitoring
If you want, I can turn this into:
- a one-page eval plan template,
- a metric/rubric spreadsheet, or
- a workflow diagram for automated + human review routing.