Prompt
I need to monitor RAG retrieval quality, chunk relevance, and citations
Latest observation
Oct 9, 2026
OpenAI APIWeb search: off
To monitor RAG retrieval quality, chunk relevance, and citations, set up evaluation at three layers:
1) Retrieval quality
Measure whether the retriever is finding the right source material.
Core metrics
- Recall@k: Did the correct source appear in the top-k results?
- MRR (Mean Reciprocal Rank): How high is the first relevant chunk ranked?
- Precision@k: How many of the retrieved chunks are actually relevant?
- Hit rate@k: Did any relevant chunk appear in top-k?
What to log
- User query
- Retrieved chunk IDs / document IDs
- Retriever scores
- Top-k ranking order
- Ground-truth relevant docs/chunks, if available
Best practice
- Build a labeled eval set of queries with known relevant sources.
- Slice results by query type, document type, and topic.
2) Chunk relevance
Measure whether each retrieved chunk is actually useful for answering the query.
Ways to assess
- Human labels: relevant / partially relevant / irrelevant
- LLM-as-judge: ask a model to score each chunk against the query
- Embedding similarity: useful for debugging, but not enough alone
Useful metrics
- Relevance rate of top-k chunks
- Average relevance score per query
- % of queries where at least one highly relevant chunk is retrieved
- Redundancy rate: how many chunks repeat the same information
Chunk-level rubric
- 2 = directly answers or contains key evidence
- 1 = somewhat related / supporting context
- 0 = irrelevant
3) Citation quality
Measure whether the final answer is properly grounded in retrieved sources.
Metrics
- Citation coverage: % of answer claims backed by citations
- Citation correctness: cited source actually supports the claim
- Citation precision: how often citations are needed vs over-cited
- Unsupported claim rate: statements with no supporting citation
- Citation diversity: are citations spread across sources appropriately?
What to check
- Each factual claim maps to one or more retrieved chunks
- Citations point to the exact chunk or source used
- No hallucinated citations
- Citations are not just present, but relevant
4) Recommended monitoring pipeline
- Log retrieval output for every query
- Run offline evaluation on a labeled benchmark set
- Use LLM-based judging for chunk relevance and answer grounding
- Track trends over time in a dashboard
- Alert on regressions after prompt, retriever, embedding, or index changes
5) Example dashboard columns
For each query:
- Query text
- Top-k retrieved chunks
- Retriever scores
- Relevant chunk found? Y/N
- Best rank of relevant chunk
- Answer citations
- Unsupported claims count
- Overall answer faithfulness score
6) Practical tooling
You can implement this with:
- RAGAS: retrieval + faithfulness-style metrics
- TruLens: groundedness, relevance, feedback functions
- LangSmith: tracing + evals
- Arize Phoenix: retrieval and embedding diagnostics
- Custom evals with an LLM judge
7) Minimal metric set to start
If you want a simple setup, start with:
- Recall@5
- MRR
- Top-5 chunk relevance rate
- Citation coverage
- Unsupported claim rate
If you want, I can also provide:
- a JSON logging schema for RAG traces,
- a Python eval script, or
- a dashboard design for these metrics.