Prompt
What should I use for LLM-as-judge workflows?
Latest observation
Oct 9, 2026
OpenAI APIWeb search: off
For LLM-as-judge workflows, a good default is:
Use a strong, instruction-following model with:
- low temperature (usually
0) - clear rubric-based prompts
- structured output (JSON / schema)
- pairwise comparisons when possible instead of absolute scoring
Model choice
Pick the best, most reliable model you can afford for the judging task. In general:
- Best quality: top-tier frontier models
- Best cost/performance: smaller but still strong reasoning models
- Avoid using a weak model to judge outputs from a stronger model unless the task is very simple
What works well in practice
1. Pairwise judging
Ask the judge to choose between two outputs:
- Which is better?
- Why?
- Return a winner and confidence
This is usually more stable than asking for a 1–10 score.
2. Rubric-based judging
Define explicit criteria, such as:
- correctness
- completeness
- helpfulness
- safety
- style
Then have the model score each dimension separately.
3. Structured outputs
Have the judge return something like:
{
"winner": "A",
"scores": {
"A": 8,
"B": 6
},
"reason": "A is more complete and accurate."
}
This makes aggregation much easier.
Good workflow tips
- Run multiple judge samples if the decision matters a lot
- Use calibration sets with human-labeled examples
- Watch for position bias and verbosity bias
- Randomize answer order
- Keep the judge blind to model identity if possible
When not to use LLM-as-judge
Avoid relying on it alone for:
- high-stakes decisions
- factual verification where tools or deterministic checks exist
- tasks with ambiguous rubrics and no human validation
Practical recommendation
If you're starting out:
- Use a strong frontier model as judge
- Use pairwise comparisons
- Force JSON output
- Randomize A/B order
- Validate against a small human-labeled set
If you want, I can also give you:
- a judge prompt template
- a JSON schema
- or a recommended evaluation pipeline for your specific use case.
Brands