Prompt
rag evaluation regression tests
Latest observation
Oct 9, 2026
OpenAI APIWeb search: off
For RAG evaluation regression tests, you’re usually trying to ensure that changes to prompts, retrieval, chunking, embeddings, reranking, or generation don’t silently make the system worse.
Here’s a practical way to set them up:
What to test
Create a fixed suite of representative queries with expected outcomes:
- Retrieval correctness
- Did the top-k results include the right source passages?
- Did relevant docs appear in the right rank range?
- Answer quality
- Is the answer factually correct?
- Does it cite the right context?
- Does it avoid hallucinating beyond retrieved evidence?
- Groundedness / faithfulness
- Is every key claim supported by retrieved context?
- Coverage
- Does the system handle common query types, edge cases, and ambiguous questions?
Recommended regression suite structure
For each test case, store:
queryexpected_answeror acceptable answer variantsexpected_sources/ doc IDsmust_includefactsmust_not_includehallucinated facts- optional metadata: domain, difficulty, intent, language
Example:
{
"id": "refund_policy_01",
"query": "What is the refund window for annual plans?",
"expected_sources": ["policy_doc_v3#section_4"],
"must_include": ["30 days"],
"must_not_include": ["60 days", "no refunds"],
"domain": "support"
}
Metrics to track
Common regression metrics:
- Recall@k / MRR / nDCG for retrieval
- Answer correctness (exact match, semantic match, or judge score)
- Faithfulness / groundedness
- Citation precision if your system cites sources
- Latency / token usage to catch performance regressions
How to automate
- Run the RAG pipeline on the fixed test set.
- Score retrieval outputs against expected sources.
- Score generated answers using:
- deterministic checks for keywords/facts
- embedding similarity
- LLM-as-judge for nuanced grading
- Compare against a baseline or thresholds.
- Fail the test if metrics drop beyond tolerance.
Good regression practices
- Keep a golden set versioned in git.
- Include both:
- happy paths
- failure modes (missing context, conflicting docs, outdated docs)
- Make tests stable: fix model version, temperature, retrieval index snapshot where possible.
- Use thresholds, not exact text matching, for generated answers.
- Add a few canary tests for critical business queries.
Example pass/fail rules
- Retrieval: relevant doc must appear in top 5
- Groundedness: at least 90% of answer claims supported by context
- Accuracy: semantic similarity above threshold or judge score >= 4/5
- No prohibited claims
Tools people often use
- RAGAS
- TruLens
- LangSmith / LangChain evals
- Promptfoo
- Custom pytest-based harnesses
If you want, I can also give you:
- a pytest template for RAG regression tests,
- a JSON schema for test cases, or
- a CI pipeline example for running them automatically.