Prompt
How do I set up regression tests for prompt changes?
Latest observation
A good way to regression-test prompt changes is to treat prompts like code: define a fixed test set, run the prompt against it, and compare outputs against expected behavior with both automatic checks and human review.
1) Define what “regression” means
Decide which failures matter:
- Format regressions: output no longer follows JSON/schema/templates
- Quality regressions: answers are less correct, less helpful, less safe
- Behavior regressions: tone changes, policy violations, missing steps
- Task regressions: the model stops performing important edge cases well
Write these down as acceptance criteria.
2) Build a stable test suite
Create a curated set of prompts that represent real usage:
- Common cases
- Edge cases
- Ambiguous inputs
- Adversarial inputs
- Known failure cases
- A few “golden” examples that must not change
For each test, store:
- Input prompt
- Expected output or expected properties
- A short rationale
- Any metadata, like category or severity
3) Prefer property-based checks over exact text matches
Exact output matching is brittle. Instead, check things like:
- Contains required fields
- Valid JSON
- No forbidden phrases
- Meets length constraints
- Mentions key facts
- Passes a rubric or classifier
- Produces one of several acceptable outputs
For creative tasks, compare against a rubric rather than exact wording.
4) Use baseline comparisons
Run the old prompt and the new prompt on the same test set and compare:
- Pass/fail rate
- Schema validity
- Human-rated quality scores
- Safety violations
- Latency and token usage
Keep a baseline snapshot so you can see whether the change improved or worsened things.
5) Automate evaluation where possible
Use a layered approach:
Deterministic checks
- Regex
- JSON schema validation
- Unit tests for prompt templates
- Required keyword checks
LLM-as-judge checks
Useful for subjective qualities:
- Correctness
- Completeness
- Tone
- Helpfulness
If you use an LLM judge, make the rubric explicit and keep it stable.
Human review
Use for:
- High-impact prompts
- Ambiguous quality criteria
- Safety-sensitive outputs
6) Control randomness
To make tests repeatable:
- Set temperature low or to 0 for regression runs
- Fix model version when possible
- Save seed if supported
- Use the same system prompt, tools, and retrieval data across runs
If you can’t fully eliminate randomness, run multiple trials and score aggregate performance.
7) Test the full prompt stack
Prompts often depend on more than the text itself:
- System prompt
- Tool definitions
- Retrieval context
- Memory/state
- Output parser
- Post-processing
Regression tests should cover the whole pipeline, not just the prompt string.
8) Store fixtures and snapshots
A practical structure:
tests/prompts/case_001.jsoncase_002.json
expected/case_001.expected.jsoncase_002.rubric.yaml
If using snapshot testing, store approved outputs and diff them when prompts change.
9) Review diffs intelligently
When outputs change, categorize the difference:
- Harmless wording change
- Better answer
- Worse answer
- Breaking format
- New unsafe behavior
This helps avoid overfitting to old wording while still catching real regressions.
10) Add CI gating
Run prompt regression tests in CI:
- On every prompt change
- Before deployment
- Nightly against a larger suite
Set thresholds like:
- Must pass all format checks
- Must not increase safety violations
- Quality score cannot drop more than X%
11) Track metrics over time
Keep a dashboard for:
- Pass rate by category
- Average rubric score
- Failure types
- Cost and latency
- Model/version used
This makes prompt changes measurable.
Simple practical setup
If you want a minimal version:
- Collect 20–100 representative examples.
- Write expected structural checks for each.
- Run old vs. new prompt on the same examples.
- Fail the test if:
- JSON/schema breaks
- Required content missing
- Safety violations increase
- Average score drops below threshold
- Manually review any changed outputs that affect important cases.
Example test case format
{
"name": "summarize_bug_report",
"input": "Summarize this bug report for engineering...",
"checks": {
"must_contain": ["root cause", "steps to reproduce"],
"must_be_valid_json": false,
"max_length": 120
},
"expected_quality": "accurate, concise, actionable"
}
Common pitfalls
- Relying only on exact string matches
- Using too few test cases
- Not freezing model/version/configuration
- Letting LLM judges drift without calibration
- Ignoring safety and format regressions
- Testing prompts in isolation when the full pipeline matters
If you want, I can also give you:
- a lightweight Python test harness,
- a CI-friendly snapshot testing setup, or
- an evaluation rubric template for LLM-as-judge regression tests.