Prompt
RAG observability citation quality tool calls
Latest observation
If you’re building RAG observability around citation quality and tool calls, here’s a practical way to think about it.
What to measure
1) Citation quality
Track whether answers are properly grounded in retrieved sources.
Useful metrics
- Citation coverage: % of factual claims with at least one citation
- Citation precision: how often cited passages actually support the claim
- Citation recall: how many supported claims are cited
- Attribution correctness: cited source matches the statement
- Source freshness: citations point to up-to-date documents
- Citation diversity: spread across sources vs. over-reliance on one chunk
Red flags
- Citations present but irrelevant
- Quotes/paraphrases that aren’t actually in the cited text
- Hallucinated citations or broken references
- All claims citing the same chunk even when multiple sources exist
2) Tool-call observability
Track how the agent uses tools and whether tool outputs improve the answer.
Useful metrics
- Tool invocation rate: how often a tool is used
- Tool success/failure rate
- Latency per tool
- Retry count
- Tool-to-answer relevance: did the tool output influence the final response?
- Tool call correctness: right tool chosen for the task
- Argument quality: were inputs to the tool valid and complete?
Red flags
- Tool called unnecessarily
- Wrong tool selected
- Repeated calls with same inputs
- Tool output ignored in final answer
- Tool call fails silently and model improvises
End-to-end RAG trace to capture
For each request, log:
- User query
- Query rewrite / decomposition
- Retrieval results
- doc IDs, chunk IDs, scores
- Reranking results
- Prompt context actually sent
- Tool calls
- tool name, args, timestamps, result/error
- Generated answer
- Citations attached to each sentence/claim
- Post-hoc evaluation
- groundedness, faithfulness, citation correctness
Recommended evaluation approach
Offline checks
Run a dataset of Q/A pairs and score:
- Answer correctness
- Groundedness
- Citation faithfulness
- Tool accuracy
Online checks
In production, sample traces and compute:
- citation errors per 100 answers
- unsupported claim rate
- tool failure rate
- mean time to answer
- retrieval-to-answer usefulness
Practical implementation tips
For citations
Store structured metadata per chunk:
doc_idchunk_idspan_start,span_endsource_urltimestampembedding_scorererank_score
When generating, require the model to emit claims with citation tags, e.g.
"The policy renews every 12 months." [doc_12#chunk_4]
Then validate:
- cited chunk contains relevant keywords/semantics
- cited chunk overlaps with the statement
- unsupported claims are flagged
For tool calls
Log each call as:
{
"tool_name": "search_docs",
"arguments": {"query": "refund policy"},
"start_time": "...",
"end_time": "...",
"status": "success",
"result_ids": ["doc_7", "doc_9"]
}
Then evaluate:
- was
search_docsactually needed? - did retrieved docs improve citation quality?
- did the final answer reflect the tool result?
Dashboards worth building
- Citation error rate over time
- Unsupported claim heatmap by topic
- Top failing retrieval queries
- Tool latency / failure rate
- Answers with no citations
- Answers where citations don’t support claims
- Cases where tools were used but not cited
If you want a simple scorecard
You can score each answer on:
- Groundedness: 0–5
- Citation correctness: 0–5
- Tool usage appropriateness: 0–5
- Tool outcome utilization: 0–5
Then combine into: RAG quality score = 40% groundedness + 30% citation correctness + 20% tool appropriateness + 10% latency/reliability
If you want, I can also give you:
- a schema for logging RAG traces,
- a SQL/JSON format for citation evaluation, or
- a LangSmith / OpenTelemetry style observability design.
Brands