Prompt
I'm building internal tools for LLM evals and need regression testing for prompts
Latest observation
Oct 9, 2026
OpenAI APIWeb search: off
A good regression-testing setup for prompts usually has 4 parts:
-
Versioned prompt artifacts
- Store prompts in git as files, not just in code strings.
- Include metadata: prompt name, version, model, temperature, system/user template, expected output schema.
-
Golden test cases
- Keep a dataset of input cases with expected properties.
- For each case, define:
- input variables
- expected structured output (if possible)
- allowed variance
- failure criteria
- Prefer “assertions on properties” over exact text matches when outputs are creative.
-
Automated eval runner
- For every prompt change, run the prompt against the golden set.
- Compare against baseline and thresholds:
- exact match / structured match
- regex / field presence
- semantic similarity
- rubric scoring via another model
- Track pass rate, deltas, and regressions by case.
-
CI gating + review workflow
- Make the eval run in CI on PRs.
- Block merges if critical tests fail or if overall score drops below threshold.
- Save artifacts: outputs, scores, diffs, logs, and prompt version used.
Practical recommendations
-
Use structured outputs whenever possible:
- JSON schema, function calling, XML tags, or constrained decoding.
- This makes regression tests much more stable.
-
Test multiple dimensions
- correctness
- completeness
- tone/style
- safety/compliance
- formatting
- latency/cost
-
Include adversarial cases
- ambiguous input
- empty input
- long input
- malformed input
- prompt injection attempts
- edge cases from production failures
-
Track non-determinism
- run each case multiple times for stochastic prompts
- compare distributions or majority outcomes, not just one sample
-
Keep a “blessed baseline”
- store a known-good output set for a specific prompt/model version
- when changing prompts, compare against that baseline
Minimal schema for a test case
{
"id": "refund_policy_001",
"input": {
"customer_message": "I was charged twice..."
},
"expected": {
"must_contain": ["refund"],
"must_not_contain": ["can't help"],
"json_schema": "support_response_v1"
},
"severity": "high"
}
Common stack choices
- Frameworks
- promptfoo
- OpenAI Evals
- LangSmith / LangGraph evals
- TruLens
- DeepEval
- Storage
- git for test cases and prompts
- sqlite/postgres for runs and results
- object storage for raw traces
Suggested workflow
- Author prompt in repo
- Add/modify golden cases
- Run local eval
- Review failures
- Merge prompt only if tests pass
- Periodic production replay to add new regression cases
If you want, I can also give you:
- a concrete repo structure for prompt regression tests,
- a Python implementation of a simple eval runner, or
- a CI pipeline example for GitHub Actions.