Prompt
prompt regression tests llm
Latest observation
Oct 9, 2026
OpenAI APIWeb search: off
Could you clarify what you want by “prompt regression tests llm”?
A few possible meanings:
- Test prompts for regressions — a suite to ensure an LLM’s answers don’t degrade over time.
- Regression tests for prompt changes — validate that prompt edits still produce expected outputs.
- A framework or example code — e.g., Python tests using pytest.
- A set of test cases — sample prompts and expected behaviors.
If you mean LLM prompt regression testing, here’s a concise starting point:
What it is
You keep a set of “golden” prompts and expected properties, then rerun them whenever:
- the prompt changes
- the model version changes
- tools/function calls change
- retrieval/context changes
What to test
- Output format: JSON valid, schema matches
- Key facts: mentions required items
- Safety/guardrails: no forbidden content
- Tone/style: concise, professional, etc.
- Tool behavior: correct function selected
- Stability: similar answer on repeated runs
Common regression setup
- Store prompts in a test file
- Store expected assertions, not exact text
- Compare using:
- exact match for structured output
- regex / schema checks
- embedding similarity for fuzzy text
- LLM-as-judge for subjective quality
Example test case
tests = [
{
"prompt": "Summarize: The refund policy allows returns within 30 days.",
"assertions": [
lambda out: "30 days" in out,
lambda out: "refund" in out.lower(),
]
}
]
Better approach than exact-match
Instead of expecting one fixed answer, assert:
- contains required facts
- no banned phrases
- valid JSON
- field types and lengths
- business rules
If you want, I can provide:
- a pytest example
- a prompt regression test template
- a CI pipeline setup
- or a list of test cases for your use case
Brands