Prompt

I'm building a RAG app, what should I log to debug bad answers later?

Technology / Observability1 observationLast seen Oct 9, 2026

Latest observation

Oct 9, 2026OpenAI APIWeb search: off

For a RAG app, the goal of logging is to let you reconstruct what the system knew, retrieved, and generated for any answer later. The best logs are the ones that help you answer:

  1. What question did the user ask?
  2. What context did retrieval return?
  3. What prompt did the model actually see?
  4. What answer was produced, with what settings?
  5. What failed: retrieval, reranking, prompt, model, or post-processing?

Log at 5 layers

1) Request / user interaction

Log enough to identify the event and reproduce it:

  • request_id / trace_id
  • session_id, conversation_id
  • user_id or anonymized stable identifier
  • timestamp
  • raw user query
  • locale / language
  • channel/app version/client type
  • conversation history used as input, or at least the messages included
  • any user feedback signals:
    • thumbs up/down
    • “bad answer” reason
    • edits, retries, follow-up questions

2) Query processing

If you do any query rewriting or classification, log:

  • original query
  • rewritten query
  • query classification / intent
  • extracted entities / filters
  • generated search filters
  • whether the query was broadened, narrowed, or decomposed into subqueries

This is useful because many bad answers come from a bad rewrite, not bad retrieval.

3) Retrieval stage

This is usually the most important debugging area.

Log:

  • retrieval method:
    • vector search / BM25 / hybrid
    • index name / namespace / shard
    • embedding model name and version
  • chunking settings used to create the index:
    • chunk size
    • overlap
    • splitter type
  • search parameters:
    • top_k
    • score threshold
    • filters applied
  • retrieved items for each hit:
    • document ID
    • chunk ID
    • source URI / title
    • score / rank
    • retrieval type if hybrid
    • chunk text or a stable reference to it
    • document metadata used for filtering
  • reranking stage:
    • reranker model/version
    • reranked scores/order
  • if retrieval failed or returned little:
    • number of candidate chunks found
    • number passed threshold
    • empty result reason if known

If you can only store one thing from retrieval, store the exact retrieved chunk IDs and text.

4) Prompt / generation input

You want to reconstruct the exact prompt that went to the model:

  • system prompt version
  • developer prompt version
  • final assembled prompt template
  • retrieved context inserted
  • citations/instructions included
  • conversation history included
  • tools/functions provided to the model
  • model name and version
  • decoding settings:
    • temperature
    • top_p
    • max_tokens
    • presence/frequency penalties
    • stop sequences
  • whether streaming was used
  • tool call traces if tools are involved:
    • tool name
    • arguments
    • tool result
    • tool errors/timeouts

5) Output / post-processing

Log:

  • raw model output
  • final user-visible answer
  • post-processing steps:
    • citation attachment
    • formatting
    • safety filtering
    • answer compression/summarization
    • language translation
  • citations actually shown to user
  • any truncation or refusal reason
  • confidence score if you generate one

Strongly recommended: log artifacts by version

A lot of debugging becomes impossible if you don’t know what changed.

Version and log:

  • prompt template version
  • embedding model version
  • reranker version
  • LLM version
  • chunking pipeline version
  • index build version
  • corpus snapshot / document version
  • retrieval configuration version

Then when an answer is bad, you can tell whether it was caused by:

  • outdated indexed content
  • changed chunking
  • different embedding model
  • a prompt tweak
  • a model upgrade

Store embeddings and corpus metadata carefully

You usually don’t need to log raw embeddings for every request, but you should log:

  • embedding model/version used
  • document corpus version
  • document ingestion timestamp
  • source doc hash / checksum
  • chunk hash

This lets you verify whether retrieval used stale or corrupted content.


Add observability metrics too

Beyond per-request logs, aggregate metrics help identify patterns:

  • retrieval hit rate
  • empty-result rate
  • average top-1 / top-5 retrieval score
  • reranker acceptance rate
  • answer citation coverage
  • hallucination/unsupported answer rate
  • latency by stage:
    • query rewrite
    • retrieval
    • rerank
    • generation
  • token usage and cost
  • fallback rates:
    • no context fallback
    • smaller model fallback
    • no-answer fallback

Practical logging structure

A useful trace often looks like:

{
  "request_id": "abc123",
  "timestamp": "2026-10-09T12:34:56Z",
  "user_id": "u_42",
  "query": {
    "original": "What are the refund terms?",
    "rewritten": "refund terms policy"
  },
  "retrieval": {
    "index_version": "kb_2026_10_01",
    "embed_model": "text-embedding-3-large",
    "top_k": 5,
    "results": [
      {
        "doc_id": "policy_17",
        "chunk_id": "policy_17_03",
        "score": 0.82,
        "text": "Refunds are available within 30 days..."
      }
    ]
  },
  "generation": {
    "model": "gpt-4.1",
    "temperature": 0.2,
    "prompt_version": "rag_prompt_v7",
    "system_prompt_version": "sys_v3"
  },
  "output": {
    "answer": "Refunds are available within 30 days ...",
    "citations": ["policy_17_03"]
  },
  "feedback": {
    "thumbs": "down"
  }
}

Privacy and safety notes

Be careful not to over-log sensitive data:

  • redact secrets, passwords, API keys
  • mask PII where possible
  • consider hashing user identifiers
  • define retention policies
  • encrypt logs at rest and in transit
  • restrict access to raw prompts and retrieved content

If your app handles regulated or sensitive data, log only what’s needed for debugging and keep the rest ephemeral.


If you want a minimal “must have” list

At minimum, log:

  • user query
  • conversation history used
  • query rewrite
  • retrieved chunk IDs + texts + scores
  • prompt template version
  • exact prompt sent to the model
  • model/version + decoding params
  • raw output
  • final output
  • citations
  • timestamps/latencies
  • feedback label

If you want, I can also give you:

  1. a production-ready JSON schema for RAG tracing, or
  2. a “what to log vs what not to log” checklist for privacy-safe observability.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.