Prompt
need llm evaluation with human review and automated scoring
Latest observation
Yes — a good LLM evaluation setup usually combines automated scoring with human review so you get both scale and judgment.
Recommended evaluation approach
1) Define what “good” means
Start by listing the dimensions you care about, for example:
- Correctness
- Completeness
- Relevance
- Factuality / hallucination rate
- Safety / policy compliance
- Tone / usefulness
- Instruction-following
- Latency / cost if needed
2) Build an evaluation set
Create a dataset of prompts representing real use cases:
- Easy / medium / hard examples
- Edge cases
- Ambiguous queries
- Adversarial or unsafe inputs
- Domain-specific questions
Include:
- Prompt
- Reference answer if available
- Metadata like category, difficulty, expected format
Automated scoring methods
Automated scoring works best when the output can be measured consistently.
Common automated metrics
- Exact match / substring match for structured outputs
- ROUGE / BLEU for summarization or translation
- F1 / precision / recall for extraction tasks
- JSON schema validation for structured responses
- Rule-based checks for formatting, citations, safety keywords, etc.
LLM-as-judge
Use a separate model to score outputs on rubrics like:
- 1–5 correctness
- 1–5 helpfulness
- Pass/fail safety
- Pairwise comparison between two model outputs
This is useful when there is no single ground truth answer, but you should:
- Use a clear rubric
- Calibrate with human-labeled examples
- Watch for judge bias
- Keep a human audit sample
Human review
Human evaluation is best for:
- Nuance
- Factual disputes
- Subjective quality
- Safety judgments
- Borderline cases
Human review process
Create a rubric with labels like:
- Pass / Partial / Fail or
- 1–5 scale for each dimension
Have reviewers assess:
- Is the answer correct?
- Did it fully address the request?
- Any hallucinations?
- Is the tone appropriate?
- Any policy issues?
Best practices
- Use 2+ reviewers for a subset to measure agreement
- Blind reviewers to model identity
- Sample difficult and random cases
- Resolve disagreements with an adjudicator
A practical combined workflow
Phase 1: Automated filter
Run all outputs through automated checks:
- Format validation
- Keyword or policy checks
- Factual consistency heuristics
- LLM judge scoring
Phase 2: Human review on a sample
Review:
- Low-scoring outputs
- Random sample of high-scoring outputs
- Edge cases
- Disagreement cases
Phase 3: Aggregate results
Track:
- Mean score by category
- Pass rate
- Error types
- Human agreement rate
- Regression over time
Suggested evaluation rubric example
For each response, score 1–5 on:
- Correctness
- Completeness
- Clarity
- Safety
- Format compliance
And store a final verdict:
- Accept
- Needs improvement
- Reject
Tooling options
Depending on your stack, you can implement this with:
- Python + pandas
- OpenAI Evals / custom scripts
- LangSmith
- Ragas for RAG systems
- DeepEval
- TruLens
- Promptfoo
- Weights & Biases for tracking
Recommended metrics to report
A strong eval report often includes:
- Overall score
- Score by task type
- Human pass rate
- Automated pass rate
- Correlation between automated and human scores
- Inter-rater agreement
- Common failure modes
- Examples of best/worst outputs
If you want a minimal setup
Use this:
- Create 100–500 representative prompts
- Generate outputs from each model/version
- Score automatically with rules + LLM judge
- Human-review 10–20% of samples
- Compare metrics and inspect failures
If you want, I can help you with any of these next:
- a human review rubric template
- an automated scoring pipeline
- a Python eval script
- an LLM-as-judge prompt
- a full evaluation framework for your use case