Prompt

How do I build regression tests for prompts in CI?

Artificial Intelligence / AI Developer Tools1 observationLast seen Oct 3, 2026

Latest observation

Oct 3, 2026OpenAI APIWeb search: off

A good way to think about prompt regression tests is: you’re not just testing text output, you’re testing that a prompt still produces the same behavior for a fixed set of representative inputs after changes to the prompt, model, tools, or surrounding code.

What to test

Use a mix of these:

  1. Golden outputs

    • Fixed input → expected output
    • Best for deterministic or near-deterministic prompts
    • Compare against saved “golden” responses
  2. Schema / contract tests

    • Output must parse as valid JSON, match a schema, include required fields, etc.
    • Great for structured outputs
  3. Semantic checks

    • Output doesn’t have to match exactly, but must satisfy rules:
      • contains key facts
      • tone is correct
      • no forbidden content
      • answer is grounded in provided context
  4. Behavioral tests

    • Multi-step tasks, tool use, classification, routing, refusal behavior
    • Example: “If user asks for a refund, route to billing tag”
  5. Non-functional tests

    • latency under threshold
    • token usage under threshold
    • no unexpected tool calls

A practical CI setup

1) Build a small eval set

Create a handful of representative cases:

  • normal inputs
  • edge cases
  • adversarial inputs
  • high-value business cases
  • known historical bugs

Example:

[
  {
    "name": "short_support_request",
    "input": "I was charged twice for my subscription.",
    "expected_contains": ["sorry", "refund", "billing"]
  },
  {
    "name": "json_output",
    "input": "Extract name and email from: Jane Doe <jane@example.com>",
    "schema": {
      "type": "object",
      "required": ["name", "email"]
    }
  }
]

2) Make outputs as deterministic as possible

In CI:

  • set temperature=0
  • set fixed model/version if possible
  • pin prompt templates
  • avoid tools/network unless tested separately

This reduces noise and makes regressions easier to detect.

3) Save baselines

For golden tests:

  • store expected outputs in version control
  • or generate approved baselines and diff against them

If exact text is too brittle, compare:

  • JSON structure
  • normalized text
  • key fields
  • embeddings/similarity for semantic checks

4) Use assertions tailored to the task

Examples:

Exact match

assert actual == expected

Contains key phrases

for phrase in expected_contains:
    assert phrase.lower() in actual.lower()

JSON schema validation

jsonschema.validate(instance=actual_json, schema=schema)

Semantic grading

  • Use a rule-based checker first
  • Optionally use an LLM-as-judge for complex cases, but keep it bounded and audited

5) Add tolerances and tiers

Don’t make every test a hard exact match. Use tiers:

  • must pass: schema valid, refusal behavior, no hallucinated tool use
  • should pass: key facts, style, completeness
  • nice to have: phrasing, formatting

This avoids flaky CI.


Recommended architecture

Minimal

  • A test file with inputs + expected checks
  • A script that calls the model
  • CI job that runs the script
  • Fail the build on regression

Better

  • Separate:
    • prompt tests
    • tool tests
    • retrieval tests
    • safety tests
  • Track metrics:
    • pass rate
    • per-case failures
    • latency
    • token usage

Best

  • Use an eval harness and store results over time
  • Run:
    • smoke tests on every PR
    • full eval suite nightly
    • production-like canary checks before release

Example CI flow

  1. PR changes prompt
  2. CI runs prompt eval suite
  3. For each test case:
    • call model with fixed settings
    • validate schema/assertions
    • compare against baseline
  4. If failures exceed threshold, block merge
  5. Optionally post a diff of changed outputs in the PR

Tips to reduce flakiness

  • Use fixed seeds if the API supports them
  • Keep temperature low
  • Limit output length
  • Normalize whitespace/punctuation before comparing
  • Avoid asserting on exact wording unless necessary
  • Mock tools and external services when testing prompt logic
  • Version your prompts and models explicitly

What not to do

  • Don’t rely only on exact string matching for all prompts
  • Don’t use a huge brittle golden set with thousands of exact outputs
  • Don’t let the model evaluate itself without checks
  • Don’t run tests with high randomness in CI

A simple strategy that works well

If you want something practical and robust:

  1. Start with 20–50 representative test cases
  2. For each case, define:
    • input
    • expected schema or key assertions
    • optional expected output example
  3. Run with temperature=0
  4. Fail on:
    • invalid schema
    • missing required facts
    • forbidden content
    • major behavioral drift
  5. Review borderline cases manually, then promote them into assertions

If you want, I can also give you:

  • a Python pytest example
  • a GitHub Actions CI example
  • a JSON schema-based prompt test harness
  • or a LangSmith / OpenAI evals style setup

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.