Prompt
I need observability for a RAG app with citation tracking and retrieval diagnostics
Latest observation
Oct 9, 2026
OpenAI APIWeb search: off
For a RAG app, observability should let you answer three questions quickly:
- Did retrieval find the right evidence?
- Did the model use that evidence correctly?
- Can I reproduce and debug failures end-to-end?
A good setup usually includes trace-level logging, retrieval metrics, citation/provenance tracking, and evaluation dashboards.
1) What to instrument
A. Request / trace metadata
For each user query, log:
trace_id,span_id- timestamp
- user/session ID
- app version, prompt version, index version
- model name, temperature, top_p
- tenant/environment
B. Retrieval diagnostics
Log for every retrieval step:
- query text
- rewritten query, if any
- embedding model used
- index/collection name
- top-k results with:
- document ID
- chunk ID
- source URI/path
- similarity score
- rank
- chunk text preview
- metadata filters applied
- retrieval latency
- whether reranker was used and its scores
- whether context compression/summarization was applied
C. Generation diagnostics
Log:
- full prompt template version
- retrieved context passed to the model
- final answer
- token usage
- generation latency
- model output confidence if available
- refusal / fallback reason
D. Citation tracking
For each citation in the answer, store:
- citation ID
- quoted/claimed span in answer
- source document ID / chunk ID
- exact supporting text span
- source URI
- retrieval rank/score at time of use
- whether citation was direct quote, paraphrase, or inferred
- coverage status:
- supported
- partially supported
- unsupported
2) Core observability features you want
A. End-to-end traces
Use tracing so a single query shows:
- user input
- retrieval pipeline
- rerank
- prompt assembly
- model call
- citations emitted
This is the easiest way to debug “bad answer, wrong source, or missing citation.”
B. Retrieval quality dashboards
Track:
- top-k hit rate
- mean reciprocal rank (MRR)
- nDCG
- recall@k
- reranker lift
- average similarity score
- empty retrieval rate
- query rewrite success rate
- filter rejection rate
C. Citation quality dashboards
Track:
- citation coverage rate
- unsupported citation rate
- answer support rate
- citation density per answer
- average citations per claim
- hallucination rate on uncited claims
- quote-to-paraphrase ratio
D. Latency/cost dashboards
Track:
- retrieval latency
- rerank latency
- model latency
- total request latency
- token counts
- cost per request
- cache hit rate
3) How to implement citation tracking
A practical pattern:
During retrieval
Assign each retrieved chunk a stable identity:
{
"doc_id": "policy_123",
"chunk_id": "policy_123#chunk_04",
"source_uri": "s3://docs/policy.pdf",
"chunk_text": "...",
"score": 0.87
}
During prompt construction
Pass a structured context block with IDs:
[1] policy_123#chunk_04
Text: Leave balances carry over up to 40 hours...
Source: s3://docs/policy.pdf
During generation
Require the model to cite using those IDs:
Employees may carry over unused leave up to 40 hours [1].
Post-generation validation
Verify each citation:
- cited chunk exists in retrieved set
- cited chunk text semantically supports the claim
- if possible, run an entailment or support-check model
- flag unsupported claims
4) Retrieval diagnostics that help in practice
Here are the most useful debugging signals:
- No relevant docs retrieved
- query too vague
- embeddings poor
- index stale
- Relevant docs retrieved but not used
- reranker misordered
- context window too small
- prompt instructions weak
- Wrong doc cited
- chunking too coarse
- duplicate or near-duplicate docs
- source metadata ambiguous
- Answer supported by retrieval but citation missing
- generation format issue
- citation parser problem
- Citation present but unsupported
- hallucination
- overcompression of context
- weak grounding instruction
5) Recommended schema
A minimal event schema:
{
"trace_id": "abc123",
"query": "What is the leave carryover policy?",
"retrieval": {
"index": "hr_docs_v4",
"top_k": [
{
"doc_id": "policy_123",
"chunk_id": "policy_123#chunk_04",
"score": 0.87,
"source_uri": "s3://docs/policy.pdf"
}
]
},
"generation": {
"model": "gpt-4.1",
"prompt_version": "rag_prompt_v7",
"answer": "Employees may carry over unused leave up to 40 hours [1]."
},
"citations": [
{
"citation_id": "1",
"chunk_id": "policy_123#chunk_04",
"status": "supported"
}
],
"metrics": {
"retrieval_latency_ms": 42,
"generation_latency_ms": 820,
"input_tokens": 1800,
"output_tokens": 112
}
}
6) Tooling options
Open-source / self-hosted
- OpenTelemetry for traces and metrics
- LangSmith for LLM/RAG tracing and evaluations
- Arize Phoenix for RAG evaluation and observability
- Weights & Biases for experiment tracking
- Prometheus + Grafana for operational dashboards
- Elastic/OpenSearch for log search
If you want a simple stack
- OpenTelemetry for traces
- Postgres or ClickHouse for event storage
- Grafana for dashboards
- Phoenix or LangSmith for RAG-specific analysis
7) Best practices
- Use stable chunk IDs so citations are reproducible.
- Log exact retrieved context used in the prompt.
- Version everything: docs, index, embeddings, prompts, rerankers.
- Separate retrieval failure from generation failure in analysis.
- Add groundedness checks for citations.
- Keep a human review UI for flagged unsupported answers.
- Sample and inspect traces regularly; dashboards alone won’t catch all issues.
8) If you want a minimal implementation plan
- Add request-level tracing with
trace_id. - Log retrieval results with chunk IDs and scores.
- Pass chunk IDs into the prompt.
- Force the model to cite chunk IDs in the answer.
- Validate citations after generation.
- Store traces in a queryable backend.
- Build dashboards for retrieval quality, citation support, latency, and cost.
If you want, I can also give you:
- a reference architecture
- a JSON schema for traces
- a Python/OpenTelemetry implementation example
- or a LangSmith/Phoenix setup for RAG citation tracking