Prompt
I'm building a way to compare LLM outputs across models and temperature settings
Latest observation
A good comparison setup usually has 3 parts:
- A fixed prompt set
- A scoring/evaluation method
- A results table that separates model vs. sampling effects
Here’s a practical way to structure it.
1) Build a benchmark prompt set
Use a mix of prompt types so you can see where models differ:
- Factual recall: short direct answers
- Reasoning: math, logic, step-by-step problems
- Instruction following: formatting, constraints, multi-part tasks
- Creative generation: writing, ideation, style imitation
- Long-context: prompts with lots of details
- Ambiguous prompts: where temperature may matter more
For each prompt, define:
- prompt ID
- category
- expected output format
- optional reference answer
- evaluation notes
2) Run each model under controlled settings
Keep everything constant except the variable you’re testing.
For model comparisons:
- same prompt
- same system message
- same max tokens
- same stop conditions
- same tool access
- same temperature/top_p settings
For temperature comparisons:
- same model
- same prompt set
- vary only temperature
- ideally run multiple samples per prompt per temperature
A useful grid:
- Models: A, B, C
- Temperatures: 0.0, 0.2, 0.7, 1.0
- Samples per cell: 3–10
3) Decide how to score outputs
You can combine automated and human evaluation.
Automated metrics
Depending on task:
- Exact match for structured answers
- F1 / ROUGE / BLEU for overlap-based tasks
- Embedding similarity for semantic closeness
- JSON/schema validity for structured outputs
- Constraint checks: length, required fields, banned terms
LLM-as-judge
Useful when outputs are open-ended:
- ask a judge model to compare two outputs
- use a rubric
- randomize order to reduce position bias
Rubric dimensions:
- correctness
- completeness
- clarity
- following instructions
- conciseness
- style/tone
Human evaluation
Best for high-stakes or nuanced tasks. Use pairwise comparisons rather than absolute scores when possible.
4) Track variance from temperature
Temperature affects output diversity, so don’t rely on one sample.
For each prompt/model/temp:
- run multiple generations
- report mean score
- report variance/std dev
- optionally report win rate over samples
This helps distinguish:
- model quality
- sampling randomness
- prompt sensitivity
5) Organize your results table
A simple schema:
| prompt_id | category | model | temperature | sample_id | output | score | judge_reason |
|---|
Then aggregate:
| model | temperature | avg_score | std_dev | win_rate |
|---|
6) Compare outputs fairly
If you’re doing pairwise comparisons:
- randomize which output appears first
- blind the judge to model name
- compare across the same prompt and temperature
- use multiple judges or repeated judgments if possible
7) Useful plots
A few that help a lot:
- score vs. temperature for each model
- box plots of scores by model/temp
- win-rate heatmap
- diversity metrics vs. temperature
- cost vs. quality scatter plot
8) Watch out for common pitfalls
- using prompts that are too easy or too hard
- mixing prompt formats
- changing max tokens between runs
- not fixing the random seed where possible
- judging outputs with hidden formatting differences
- relying on one sample at nonzero temperature
9) Minimal experimental design
If you want a simple but solid setup:
- 50 prompts
- 3 models
- 4 temperatures
- 5 samples each
- human or judge-based pairwise scoring
- report average score and confidence intervals
That gives: 50 × 3 × 4 × 5 = 3,000 generations, which is enough to see meaningful patterns.
If you want, I can help you design:
- a CSV/JSON schema for storing results,
- a Python evaluation pipeline, or
- a scoring rubric for your specific use case.