Prompt
How do I compare outputs across two prompt versions before a rollout?
Latest observation
A solid way to compare two prompt versions before rollout is to do a side-by-side evaluation on the same inputs, then score both automatically and manually.
Practical workflow
1) Freeze a test set
Use a representative set of prompts:
- common cases
- edge cases
- failure cases
- a few “golden” examples with known good outputs
Keep the inputs identical for both versions.
2) Run both prompts against the same inputs
For each test case, generate:
- output from prompt v1
- output from prompt v2
Use the same model and settings where possible:
- temperature
- max tokens
- tools/function calls
- retrieval context
If you’re testing stochastic behavior, run each case multiple times.
3) Compare with clear criteria
Decide upfront what “better” means. Common criteria:
- correctness
- completeness
- formatting/structure
- tone/style
- safety/compliance
- latency/cost
- robustness on edge cases
4) Use a scoring rubric
A simple rubric works well, for example:
- 2 = better
- 1 = same
- 0 = worse
Or score each dimension from 1–5. Keep the rubric specific so different reviewers agree.
5) Review outputs blind if possible
Hide which prompt version produced which output to reduce bias. Then have reviewers choose:
- A wins
- B wins
- tie
- both bad
6) Aggregate results
Look at:
- win rate per version
- average score by dimension
- failure rate on critical cases
- regression categories
If prompt v2 improves most cases but fails a few critical ones, that matters more than the average.
7) Check cost and latency
Don’t just compare quality:
- token usage
- response time
- tool-call frequency
- retry rate
A prompt that is slightly better but much slower may not be worth it.
Good tools / methods
Depending on your stack, you can use:
- pairwise human evals
- LLM-as-judge for first-pass scoring
- automated tests for schema/format validity
- A/B tests in staging or limited production
- offline regression suites for repeatability
Recommended rollout approach
- Offline compare v1 vs v2 on a fixed eval set
- Fix obvious regressions
- Run a small canary or shadow test
- Roll out gradually with monitoring
Minimal template for comparison
For each test case:
| Input | v1 output | v2 output | Winner | Reason |
|---|---|---|---|---|
| ... | ... | ... | ... | ... |
Key tip
Use the same test set and same evaluation rubric every time. That gives you a reliable regression baseline across prompt versions.
If you want, I can also give you:
- a prompt comparison spreadsheet template
- a Python script for pairwise evals
- or an LLM-judge rubric for your specific use case.