Prompt
I'm building an LLM application and need continuous evals in GitHub Actions
Latest observation
Oct 9, 2026
OpenAI APIWeb search: off
Here’s a practical way to set up continuous evals in GitHub Actions for an LLM app.
Recommended approach
Treat evals like tests:
- Run on every PR for a small, fast smoke suite
- Run nightly or on merge for a larger regression suite
- Fail the workflow if quality drops below a threshold
- Store results as artifacts so you can compare over time
Typical architecture
1) Define eval cases
Keep a versioned dataset in your repo, for example:
evals/
prompts.jsonl
expected.jsonl
Each row might include:
- prompt
- reference answer or rubric
- expected tool calls / structured output
- tags like
billing,retrieval,safety
2) Write an eval runner
Use Python, Node, or whatever your app uses. The runner should:
- call your app or model endpoint
- score outputs
- print a summary
- exit non-zero if thresholds are not met
Example metrics:
- exact match
- semantic similarity
- JSON/schema validity
- tool-call correctness
- hallucination / groundedness checks
- human rubric if you have an LLM-as-judge step
3) Add GitHub Actions workflow
A simple workflow might:
- checkout code
- install dependencies
- set secrets for model/API access
- run evals
- upload report artifact
- optionally comment on PR with results
Example GitHub Actions workflow
name: llm-evals
on:
pull_request:
push:
branches: [main]
schedule:
- cron: "0 3 * * *" # nightly at 3 AM UTC
jobs:
evals:
runs-on: ubuntu-latest
steps:
- name: Checkout
uses: actions/checkout@v4
- name: Set up Python
uses: actions/setup-python@v5
with:
python-version: "3.11"
- name: Install dependencies
run: |
pip install -r requirements.txt
- name: Run evals
env:
OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
OTHER_API_KEY: ${{ secrets.OTHER_API_KEY }}
run: |
python -m evals.run --suite smoke --output results.json
- name: Upload eval results
uses: actions/upload-artifact@v4
with:
name: llm-eval-results
path: results.json
Example eval runner pattern
# evals/run.py
import json
import sys
THRESHOLD = 0.85
def run_suite(suite_name: str):
# Load test cases, call your app, score outputs
# Return a dict like:
return {
"suite": suite_name,
"score": 0.91,
"passed": True,
"details": [
{"id": "case1", "score": 1.0},
{"id": "case2", "score": 0.82},
],
}
if __name__ == "__main__":
import argparse
parser = argparse.ArgumentParser()
parser.add_argument("--suite", default="smoke")
parser.add_argument("--output", default="results.json")
args = parser.parse_args()
result = run_suite(args.suite)
with open(args.output, "w") as f:
json.dump(result, f, indent=2)
print(json.dumps(result, indent=2))
if result["score"] < THRESHOLD:
sys.exit(1)
Good practices
Keep CI fast
- Use a smoke suite for PRs
- Limit number of cases per run
- Cache dependencies
- Avoid expensive judge calls unless necessary
Separate stable and flaky metrics
- For PR gating, use deterministic or low-variance checks
- For broader quality trends, run larger suites nightly
Version everything
Track:
- prompts
- expected outputs
- scoring logic
- model version
- retrieval corpus version
Store historical results
Use one of:
- GitHub Actions artifacts
- commit comments
- a metrics dashboard
- external tools like LangSmith, W&B, Arize, Helicone, etc.
Gate on deltas, not just absolute score
Example:
- fail if score < 0.85
- or fail if score drops by > 3% from baseline
If your app uses structured outputs
Add schema validation as an eval:
from pydantic import BaseModel, ValidationError
class Response(BaseModel):
answer: str
citations: list[str]
def validate_output(text):
try:
Response.model_validate_json(text)
return True
except ValidationError:
return False
This is especially useful for:
- tool calling
- agents
- JSON mode outputs
- extraction pipelines
If you want LLM-as-judge
Use it carefully:
- keep judge prompts versioned
- fix the judge model version if possible
- use pairwise comparisons for regressions
- sample a subset for PRs, larger set nightly
You can also combine:
- deterministic checks
- LLM judge
- human review for edge cases
Suggested workflow strategy
PR workflow
- 10–50 cases
- fast checks
- schema validation
- high-signal judge metrics
- comment summary on PR
Main branch workflow
- full regression suite
- compare to baseline
- upload artifact
- maybe update dashboard
Scheduled workflow
- all evals
- longer judge chains
- trend tracking
- alerting on regressions
If you want, I can help you with one of these next:
- A ready-to-paste GitHub Actions YAML
- A Python eval harness template
- An integration with OpenAI / LangSmith / promptfoo / pytest
- A setup for PR comments with eval summaries
If you tell me your stack, I can tailor it.