Prompt
How do I build regression tests for prompts in CI?
Latest observation
A good way to think about prompt regression tests is: you’re not just testing text output, you’re testing that a prompt still produces the same behavior for a fixed set of representative inputs after changes to the prompt, model, tools, or surrounding code.
What to test
Use a mix of these:
-
Golden outputs
- Fixed input → expected output
- Best for deterministic or near-deterministic prompts
- Compare against saved “golden” responses
-
Schema / contract tests
- Output must parse as valid JSON, match a schema, include required fields, etc.
- Great for structured outputs
-
Semantic checks
- Output doesn’t have to match exactly, but must satisfy rules:
- contains key facts
- tone is correct
- no forbidden content
- answer is grounded in provided context
- Output doesn’t have to match exactly, but must satisfy rules:
-
Behavioral tests
- Multi-step tasks, tool use, classification, routing, refusal behavior
- Example: “If user asks for a refund, route to billing tag”
-
Non-functional tests
- latency under threshold
- token usage under threshold
- no unexpected tool calls
A practical CI setup
1) Build a small eval set
Create a handful of representative cases:
- normal inputs
- edge cases
- adversarial inputs
- high-value business cases
- known historical bugs
Example:
[
{
"name": "short_support_request",
"input": "I was charged twice for my subscription.",
"expected_contains": ["sorry", "refund", "billing"]
},
{
"name": "json_output",
"input": "Extract name and email from: Jane Doe <jane@example.com>",
"schema": {
"type": "object",
"required": ["name", "email"]
}
}
]
2) Make outputs as deterministic as possible
In CI:
- set
temperature=0 - set fixed model/version if possible
- pin prompt templates
- avoid tools/network unless tested separately
This reduces noise and makes regressions easier to detect.
3) Save baselines
For golden tests:
- store expected outputs in version control
- or generate approved baselines and diff against them
If exact text is too brittle, compare:
- JSON structure
- normalized text
- key fields
- embeddings/similarity for semantic checks
4) Use assertions tailored to the task
Examples:
Exact match
assert actual == expected
Contains key phrases
for phrase in expected_contains:
assert phrase.lower() in actual.lower()
JSON schema validation
jsonschema.validate(instance=actual_json, schema=schema)
Semantic grading
- Use a rule-based checker first
- Optionally use an LLM-as-judge for complex cases, but keep it bounded and audited
5) Add tolerances and tiers
Don’t make every test a hard exact match. Use tiers:
- must pass: schema valid, refusal behavior, no hallucinated tool use
- should pass: key facts, style, completeness
- nice to have: phrasing, formatting
This avoids flaky CI.
Recommended architecture
Minimal
- A test file with inputs + expected checks
- A script that calls the model
- CI job that runs the script
- Fail the build on regression
Better
- Separate:
- prompt tests
- tool tests
- retrieval tests
- safety tests
- Track metrics:
- pass rate
- per-case failures
- latency
- token usage
Best
- Use an eval harness and store results over time
- Run:
- smoke tests on every PR
- full eval suite nightly
- production-like canary checks before release
Example CI flow
- PR changes prompt
- CI runs prompt eval suite
- For each test case:
- call model with fixed settings
- validate schema/assertions
- compare against baseline
- If failures exceed threshold, block merge
- Optionally post a diff of changed outputs in the PR
Tips to reduce flakiness
- Use fixed seeds if the API supports them
- Keep temperature low
- Limit output length
- Normalize whitespace/punctuation before comparing
- Avoid asserting on exact wording unless necessary
- Mock tools and external services when testing prompt logic
- Version your prompts and models explicitly
What not to do
- Don’t rely only on exact string matching for all prompts
- Don’t use a huge brittle golden set with thousands of exact outputs
- Don’t let the model evaluate itself without checks
- Don’t run tests with high randomness in CI
A simple strategy that works well
If you want something practical and robust:
- Start with 20–50 representative test cases
- For each case, define:
- input
- expected schema or key assertions
- optional expected output example
- Run with
temperature=0 - Fail on:
- invalid schema
- missing required facts
- forbidden content
- major behavioral drift
- Review borderline cases manually, then promote them into assertions
If you want, I can also give you:
- a Python pytest example
- a GitHub Actions CI example
- a JSON schema-based prompt test harness
- or a LangSmith / OpenAI evals style setup