Prompt
How do I run continuous evaluation for prompt changes in CI?
Latest observation
Oct 9, 2026
OpenAI APIWeb search: off
A good pattern is to treat prompts like code: version them, test them, and block merges if they regress.
Basic CI setup
-
Keep a small but meaningful eval set
- 20–200 examples is common to start.
- Include “golden” inputs that represent your real traffic.
- Cover edge cases, adversarial inputs, and expected failure modes.
-
Define measurable criteria
- Exact match / regex checks for structured outputs
- JSON schema validation
- LLM-as-judge for subjective tasks
- Task-specific scores like accuracy, pass@k, win-rate, or rubric-based grades
-
Run evals on every prompt change
- Diff the prompt file in PRs
- Execute the model against the eval set
- Compare current vs baseline results
-
Fail the build on regressions
- Set thresholds like:
- overall score must not drop by more than 2%
- no critical test case may fail
- structured output pass rate must be 100%
- Optionally allow a human review gate for borderline cases
- Set thresholds like:
Recommended CI flow
1) Store prompts and evals in repo
Example structure:
prompts/
support_agent.prompt.md
evals/
support_agent_cases.jsonl
scripts/
run_eval.py
2) Baseline the current best prompt
You can:
- save a “golden” prompt version in main
- or store baseline scores as artifacts from the last green build
3) In CI, run:
- prompt linting
- eval execution
- score comparison
- report generation
Example GitHub Actions workflow
name: Prompt Eval CI
on:
pull_request:
paths:
- "prompts/**"
- "evals/**"
- "scripts/**"
jobs:
eval:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Set up Python
uses: actions/setup-python@v5
with:
python-version: "3.11"
- name: Install deps
run: |
pip install -r requirements.txt
- name: Run prompt evals
env:
OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
run: |
python scripts/run_eval.py \
--prompt prompts/support_agent.prompt.md \
--evals evals/support_agent_cases.jsonl \
--baseline artifacts/baseline.json \
--output artifacts/current.json
- name: Upload results
uses: actions/upload-artifact@v4
with:
name: eval-results
path: artifacts/
Example eval script behavior
Your script should:
- load test cases
- render the prompt with each input
- call the model
- score each output
- aggregate results
- exit non-zero if thresholds fail
Pseudo-logic:
if current_score < baseline_score - allowed_drop:
raise SystemExit("Regression detected")
if critical_failures > 0:
raise SystemExit("Critical eval failure")
Practical scoring approach
For deterministic tasks
Use:
- exact string match
- JSON schema validation
- unit-test style assertions
For open-ended tasks
Use:
- rubric-based LLM judge
- pairwise comparison vs baseline
- human review for a sampled subset
A very common approach is:
- hard checks for format/safety
- soft checks for quality
Make it robust
-
Seed and temperature
- Set temperature low for CI to reduce noise.
- Run multiple samples if your application depends on stochasticity.
-
Use canary thresholds
- Don’t fail on tiny differences from model nondeterminism.
- Require statistically meaningful regressions before blocking.
-
Keep evals fast
- Run a small “smoke eval” on every PR
- Run a larger nightly evaluation
-
Track trends
- Save scores over time so you can detect drift.
A good CI strategy
- PR checks: quick eval on a subset
- Main branch: full eval suite
- Nightly: broader regression and drift testing
- Release gate: manual approval if scores are near threshold
Tools you can use
Depending on your stack, you can use:
- custom Python scripts
- OpenAI Evals-style harnesses
- promptfoo
- LangSmith / LangChain evals
- DeepEval
- TruLens
- Ragas for RAG workflows
Rule of thumb
Start simple:
- one prompt file
- one JSONL eval set
- one score threshold
- one CI job
Then expand into:
- multiple metrics
- LLM judges
- nightly regression suites
- model/prompt comparison dashboards
If you want, I can also give you:
- a working GitHub Actions + Python eval example
- a promptfoo CI setup
- or a template for scoring prompt regressions