Prompt

I'm building an LLM application and need continuous evals in GitHub Actions

Artificial Intelligence / AI Developer Tools1 observationLast seen Oct 9, 2026

Latest observation

Oct 9, 2026OpenAI APIWeb search: off

Here’s a practical way to set up continuous evals in GitHub Actions for an LLM app.

Recommended approach

Treat evals like tests:

  • Run on every PR for a small, fast smoke suite
  • Run nightly or on merge for a larger regression suite
  • Fail the workflow if quality drops below a threshold
  • Store results as artifacts so you can compare over time

Typical architecture

1) Define eval cases

Keep a versioned dataset in your repo, for example:

evals/
  prompts.jsonl
  expected.jsonl

Each row might include:

  • prompt
  • reference answer or rubric
  • expected tool calls / structured output
  • tags like billing, retrieval, safety

2) Write an eval runner

Use Python, Node, or whatever your app uses. The runner should:

  • call your app or model endpoint
  • score outputs
  • print a summary
  • exit non-zero if thresholds are not met

Example metrics:

  • exact match
  • semantic similarity
  • JSON/schema validity
  • tool-call correctness
  • hallucination / groundedness checks
  • human rubric if you have an LLM-as-judge step

3) Add GitHub Actions workflow

A simple workflow might:

  • checkout code
  • install dependencies
  • set secrets for model/API access
  • run evals
  • upload report artifact
  • optionally comment on PR with results

Example GitHub Actions workflow

name: llm-evals

on:
  pull_request:
  push:
    branches: [main]
  schedule:
    - cron: "0 3 * * *" # nightly at 3 AM UTC

jobs:
  evals:
    runs-on: ubuntu-latest

    steps:
      - name: Checkout
        uses: actions/checkout@v4

      - name: Set up Python
        uses: actions/setup-python@v5
        with:
          python-version: "3.11"

      - name: Install dependencies
        run: |
          pip install -r requirements.txt

      - name: Run evals
        env:
          OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
          OTHER_API_KEY: ${{ secrets.OTHER_API_KEY }}
        run: |
          python -m evals.run --suite smoke --output results.json

      - name: Upload eval results
        uses: actions/upload-artifact@v4
        with:
          name: llm-eval-results
          path: results.json

Example eval runner pattern

# evals/run.py
import json
import sys

THRESHOLD = 0.85

def run_suite(suite_name: str):
    # Load test cases, call your app, score outputs
    # Return a dict like:
    return {
        "suite": suite_name,
        "score": 0.91,
        "passed": True,
        "details": [
            {"id": "case1", "score": 1.0},
            {"id": "case2", "score": 0.82},
        ],
    }

if __name__ == "__main__":
    import argparse
    parser = argparse.ArgumentParser()
    parser.add_argument("--suite", default="smoke")
    parser.add_argument("--output", default="results.json")
    args = parser.parse_args()

    result = run_suite(args.suite)

    with open(args.output, "w") as f:
        json.dump(result, f, indent=2)

    print(json.dumps(result, indent=2))

    if result["score"] < THRESHOLD:
        sys.exit(1)

Good practices

Keep CI fast

  • Use a smoke suite for PRs
  • Limit number of cases per run
  • Cache dependencies
  • Avoid expensive judge calls unless necessary

Separate stable and flaky metrics

  • For PR gating, use deterministic or low-variance checks
  • For broader quality trends, run larger suites nightly

Version everything

Track:

  • prompts
  • expected outputs
  • scoring logic
  • model version
  • retrieval corpus version

Store historical results

Use one of:

  • GitHub Actions artifacts
  • commit comments
  • a metrics dashboard
  • external tools like LangSmith, W&B, Arize, Helicone, etc.

Gate on deltas, not just absolute score

Example:

  • fail if score < 0.85
  • or fail if score drops by > 3% from baseline

If your app uses structured outputs

Add schema validation as an eval:

from pydantic import BaseModel, ValidationError

class Response(BaseModel):
    answer: str
    citations: list[str]

def validate_output(text):
    try:
        Response.model_validate_json(text)
        return True
    except ValidationError:
        return False

This is especially useful for:

  • tool calling
  • agents
  • JSON mode outputs
  • extraction pipelines

If you want LLM-as-judge

Use it carefully:

  • keep judge prompts versioned
  • fix the judge model version if possible
  • use pairwise comparisons for regressions
  • sample a subset for PRs, larger set nightly

You can also combine:

  • deterministic checks
  • LLM judge
  • human review for edge cases

Suggested workflow strategy

PR workflow

  • 10–50 cases
  • fast checks
  • schema validation
  • high-signal judge metrics
  • comment summary on PR

Main branch workflow

  • full regression suite
  • compare to baseline
  • upload artifact
  • maybe update dashboard

Scheduled workflow

  • all evals
  • longer judge chains
  • trend tracking
  • alerting on regressions

If you want, I can help you with one of these next:

  1. A ready-to-paste GitHub Actions YAML
  2. A Python eval harness template
  3. An integration with OpenAI / LangSmith / promptfoo / pytest
  4. A setup for PR comments with eval summaries

If you tell me your stack, I can tailor it.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.