Prompt
I'm building a RAG app, what should I log to debug bad answers later?
Latest observation
For a RAG app, the goal of logging is to let you reconstruct what the system knew, retrieved, and generated for any answer later. The best logs are the ones that help you answer:
- What question did the user ask?
- What context did retrieval return?
- What prompt did the model actually see?
- What answer was produced, with what settings?
- What failed: retrieval, reranking, prompt, model, or post-processing?
Log at 5 layers
1) Request / user interaction
Log enough to identify the event and reproduce it:
request_id/trace_idsession_id,conversation_iduser_idor anonymized stable identifier- timestamp
- raw user query
- locale / language
- channel/app version/client type
- conversation history used as input, or at least the messages included
- any user feedback signals:
- thumbs up/down
- “bad answer” reason
- edits, retries, follow-up questions
2) Query processing
If you do any query rewriting or classification, log:
- original query
- rewritten query
- query classification / intent
- extracted entities / filters
- generated search filters
- whether the query was broadened, narrowed, or decomposed into subqueries
This is useful because many bad answers come from a bad rewrite, not bad retrieval.
3) Retrieval stage
This is usually the most important debugging area.
Log:
- retrieval method:
- vector search / BM25 / hybrid
- index name / namespace / shard
- embedding model name and version
- chunking settings used to create the index:
- chunk size
- overlap
- splitter type
- search parameters:
- top_k
- score threshold
- filters applied
- retrieved items for each hit:
- document ID
- chunk ID
- source URI / title
- score / rank
- retrieval type if hybrid
- chunk text or a stable reference to it
- document metadata used for filtering
- reranking stage:
- reranker model/version
- reranked scores/order
- if retrieval failed or returned little:
- number of candidate chunks found
- number passed threshold
- empty result reason if known
If you can only store one thing from retrieval, store the exact retrieved chunk IDs and text.
4) Prompt / generation input
You want to reconstruct the exact prompt that went to the model:
- system prompt version
- developer prompt version
- final assembled prompt template
- retrieved context inserted
- citations/instructions included
- conversation history included
- tools/functions provided to the model
- model name and version
- decoding settings:
- temperature
- top_p
- max_tokens
- presence/frequency penalties
- stop sequences
- whether streaming was used
- tool call traces if tools are involved:
- tool name
- arguments
- tool result
- tool errors/timeouts
5) Output / post-processing
Log:
- raw model output
- final user-visible answer
- post-processing steps:
- citation attachment
- formatting
- safety filtering
- answer compression/summarization
- language translation
- citations actually shown to user
- any truncation or refusal reason
- confidence score if you generate one
Strongly recommended: log artifacts by version
A lot of debugging becomes impossible if you don’t know what changed.
Version and log:
- prompt template version
- embedding model version
- reranker version
- LLM version
- chunking pipeline version
- index build version
- corpus snapshot / document version
- retrieval configuration version
Then when an answer is bad, you can tell whether it was caused by:
- outdated indexed content
- changed chunking
- different embedding model
- a prompt tweak
- a model upgrade
Store embeddings and corpus metadata carefully
You usually don’t need to log raw embeddings for every request, but you should log:
- embedding model/version used
- document corpus version
- document ingestion timestamp
- source doc hash / checksum
- chunk hash
This lets you verify whether retrieval used stale or corrupted content.
Add observability metrics too
Beyond per-request logs, aggregate metrics help identify patterns:
- retrieval hit rate
- empty-result rate
- average top-1 / top-5 retrieval score
- reranker acceptance rate
- answer citation coverage
- hallucination/unsupported answer rate
- latency by stage:
- query rewrite
- retrieval
- rerank
- generation
- token usage and cost
- fallback rates:
- no context fallback
- smaller model fallback
- no-answer fallback
Practical logging structure
A useful trace often looks like:
{
"request_id": "abc123",
"timestamp": "2026-10-09T12:34:56Z",
"user_id": "u_42",
"query": {
"original": "What are the refund terms?",
"rewritten": "refund terms policy"
},
"retrieval": {
"index_version": "kb_2026_10_01",
"embed_model": "text-embedding-3-large",
"top_k": 5,
"results": [
{
"doc_id": "policy_17",
"chunk_id": "policy_17_03",
"score": 0.82,
"text": "Refunds are available within 30 days..."
}
]
},
"generation": {
"model": "gpt-4.1",
"temperature": 0.2,
"prompt_version": "rag_prompt_v7",
"system_prompt_version": "sys_v3"
},
"output": {
"answer": "Refunds are available within 30 days ...",
"citations": ["policy_17_03"]
},
"feedback": {
"thumbs": "down"
}
}
Privacy and safety notes
Be careful not to over-log sensitive data:
- redact secrets, passwords, API keys
- mask PII where possible
- consider hashing user identifiers
- define retention policies
- encrypt logs at rest and in transit
- restrict access to raw prompts and retrieved content
If your app handles regulated or sensitive data, log only what’s needed for debugging and keep the rest ephemeral.
If you want a minimal “must have” list
At minimum, log:
- user query
- conversation history used
- query rewrite
- retrieved chunk IDs + texts + scores
- prompt template version
- exact prompt sent to the model
- model/version + decoding params
- raw output
- final output
- citations
- timestamps/latencies
- feedback label
If you want, I can also give you:
- a production-ready JSON schema for RAG tracing, or
- a “what to log vs what not to log” checklist for privacy-safe observability.