Prompt
need llm evaluation for open ended answers without exact match
Latest observation
For open-ended answers, exact match is usually the wrong metric. Better evaluation depends on whether you want to measure correctness, completeness, relevance, or quality of reasoning.
Common evaluation options
1. LLM-as-a-judge
Use another model to score the answer against a reference answer or rubric.
Best for: subjective/open-ended tasks like summarization, explanation, helpfulness.
How: ask the judge model to rate dimensions such as:
- correctness
- completeness
- clarity
- relevance
- faithfulness to source
Example prompt for judge:
Compare the candidate answer to the reference answer. Score from 1–5 on:
- factual correctness
- completeness
- relevance
- conciseness
Return a JSON object with scores and short justification.
2. Semantic similarity
Use embedding-based similarity instead of exact string match.
Metrics:
- cosine similarity between embeddings
- BERTScore
- BLEURT
- ROUGE-L / ROUGE-1 / ROUGE-2
Best for: answers that can be paraphrased but still mean the same thing.
Limitation: similarity can miss factual errors or reward fluent but wrong answers.
3. Rubric-based human evaluation
Create a rubric and let humans score outputs.
Example rubric:
- 0 = incorrect
- 1 = partially correct
- 2 = mostly correct
- 3 = fully correct and well explained
Best for: high-stakes evaluation and benchmarking.
Limitation: expensive and slower.
4. Task-specific automatic metrics
If the open-ended task can be decomposed, evaluate subparts.
Examples:
- QA: answer contains key facts
- summarization: coverage + factual consistency
- code generation: unit tests pass
- retrieval-grounded QA: answer supported by retrieved context
5. Pairwise comparison
Compare two model outputs and ask which is better.
Best for: ranking models when no single gold answer exists.
This is often more reliable than absolute scoring.
Recommended approach
For most open-ended LLM evaluation, use:
- A rubric
- LLM-as-judge
- Human spot checks
- Pairwise comparison if comparing models
- Embeddings / ROUGE / BERTScore as supporting signals only
Simple practical setup
If you want something easy and scalable:
- Prepare a small set of reference answers or criteria
- Use an LLM judge with a fixed rubric
- Score each answer on 1–5 for:
- correctness
- completeness
- reasoning quality
- Aggregate average scores
- Validate a sample with human review
Example judge output format
{
"correctness": 4,
"completeness": 3,
"clarity": 5,
"overall": 4,
"notes": "Mostly correct, but missed one important detail."
}
If you want, I can give you:
- a ready-to-use evaluation prompt
- a Python script
- or a full evaluation framework for your specific use case.