Prompt

How do I run continuous evaluation for prompt changes in CI?

Artificial Intelligence / AI Developer Tools1 observationLast seen Oct 9, 2026

Latest observation

Oct 9, 2026OpenAI APIWeb search: off

A good pattern is to treat prompts like code: version them, test them, and block merges if they regress.

Basic CI setup

  1. Keep a small but meaningful eval set

    • 20–200 examples is common to start.
    • Include “golden” inputs that represent your real traffic.
    • Cover edge cases, adversarial inputs, and expected failure modes.
  2. Define measurable criteria

    • Exact match / regex checks for structured outputs
    • JSON schema validation
    • LLM-as-judge for subjective tasks
    • Task-specific scores like accuracy, pass@k, win-rate, or rubric-based grades
  3. Run evals on every prompt change

    • Diff the prompt file in PRs
    • Execute the model against the eval set
    • Compare current vs baseline results
  4. Fail the build on regressions

    • Set thresholds like:
      • overall score must not drop by more than 2%
      • no critical test case may fail
      • structured output pass rate must be 100%
    • Optionally allow a human review gate for borderline cases

Recommended CI flow

1) Store prompts and evals in repo

Example structure:

prompts/
  support_agent.prompt.md
evals/
  support_agent_cases.jsonl
scripts/
  run_eval.py

2) Baseline the current best prompt

You can:

  • save a “golden” prompt version in main
  • or store baseline scores as artifacts from the last green build

3) In CI, run:

  • prompt linting
  • eval execution
  • score comparison
  • report generation

Example GitHub Actions workflow

name: Prompt Eval CI

on:
  pull_request:
    paths:
      - "prompts/**"
      - "evals/**"
      - "scripts/**"

jobs:
  eval:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4

      - name: Set up Python
        uses: actions/setup-python@v5
        with:
          python-version: "3.11"

      - name: Install deps
        run: |
          pip install -r requirements.txt

      - name: Run prompt evals
        env:
          OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
        run: |
          python scripts/run_eval.py \
            --prompt prompts/support_agent.prompt.md \
            --evals evals/support_agent_cases.jsonl \
            --baseline artifacts/baseline.json \
            --output artifacts/current.json

      - name: Upload results
        uses: actions/upload-artifact@v4
        with:
          name: eval-results
          path: artifacts/

Example eval script behavior

Your script should:

  • load test cases
  • render the prompt with each input
  • call the model
  • score each output
  • aggregate results
  • exit non-zero if thresholds fail

Pseudo-logic:

if current_score < baseline_score - allowed_drop:
    raise SystemExit("Regression detected")
if critical_failures > 0:
    raise SystemExit("Critical eval failure")

Practical scoring approach

For deterministic tasks

Use:

  • exact string match
  • JSON schema validation
  • unit-test style assertions

For open-ended tasks

Use:

  • rubric-based LLM judge
  • pairwise comparison vs baseline
  • human review for a sampled subset

A very common approach is:

  • hard checks for format/safety
  • soft checks for quality

Make it robust

  • Seed and temperature

    • Set temperature low for CI to reduce noise.
    • Run multiple samples if your application depends on stochasticity.
  • Use canary thresholds

    • Don’t fail on tiny differences from model nondeterminism.
    • Require statistically meaningful regressions before blocking.
  • Keep evals fast

    • Run a small “smoke eval” on every PR
    • Run a larger nightly evaluation
  • Track trends

    • Save scores over time so you can detect drift.

A good CI strategy

  • PR checks: quick eval on a subset
  • Main branch: full eval suite
  • Nightly: broader regression and drift testing
  • Release gate: manual approval if scores are near threshold

Tools you can use

Depending on your stack, you can use:

  • custom Python scripts
  • OpenAI Evals-style harnesses
  • promptfoo
  • LangSmith / LangChain evals
  • DeepEval
  • TruLens
  • Ragas for RAG workflows

Rule of thumb

Start simple:

  1. one prompt file
  2. one JSONL eval set
  3. one score threshold
  4. one CI job

Then expand into:

  • multiple metrics
  • LLM judges
  • nightly regression suites
  • model/prompt comparison dashboards

If you want, I can also give you:

  • a working GitHub Actions + Python eval example
  • a promptfoo CI setup
  • or a template for scoring prompt regressions

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.