Prompt

I'm building a RAG chatbot, how do I inspect citations and source chunks?

Technology / Observability1 observationLast seen Oct 9, 2026

Latest observation

Oct 9, 2026OpenAI APIWeb search: off

To inspect citations and source chunks in a RAG chatbot, the key is to make your retrieval and generation pipeline observable end-to-end.

What to log/store

For each user question, capture:

  1. User query
  2. Retrieved chunks
    • chunk text
    • document/file name
    • chunk ID
    • page/section/URL
    • retrieval score / rank
  3. Prompt sent to the LLM
  4. Model response
  5. Citations in the response
    • which answer sentence maps to which chunk/source
    • citation labels like [1], [2]
  6. Metadata
    • timestamp
    • embedding model
    • retriever config
    • reranker config
    • conversation/session ID

Best practice: assign stable source IDs

When you chunk documents, give every chunk a stable identifier, for example:

  • doc_042#chunk_003
  • employee_handbook:p12:c2
  • https://example.com/policy#section-4

Then every retrieved chunk can be referenced in the answer and in your logs.

Make citations explicit in the LLM prompt

Ask the model to cite from provided sources only. Example:

Use only the sources below. When you use a source, cite it as [1], [2], etc.
If the answer is not in the sources, say you don't know.

Sources:
[1] Chunk text...
[2] Chunk text...

This makes it easier to inspect why the model answered the way it did.

Return structured retrieval data

Instead of only returning plain text, return something like:

{
  "answer": "The policy allows ... [1]",
  "citations": [
    {
      "label": "[1]",
      "chunk_id": "employee_handbook:p12:c2",
      "document": "employee_handbook.pdf",
      "page": 12,
      "text": "..."
    }
  ],
  "retrieved_chunks": [
    {
      "chunk_id": "employee_handbook:p12:c2",
      "score": 0.87,
      "text": "..."
    }
  ]
}

How to inspect source chunks during debugging

1. Print the retrieved chunks

For every query, inspect top-k results:

  • Are they actually relevant?
  • Are chunks too large or too small?
  • Are duplicates being retrieved?

2. Compare citations to sources

Check whether each citation points to:

  • the right chunk
  • the right document
  • the right page/section

3. Trace answer spans

If you want stronger grounding, map answer sentences to source spans:

  • sentence-level attribution
  • token-level highlighting
  • quote extraction

4. Use a trace UI

Tools that help:

  • LangSmith
  • OpenTelemetry-based tracing
  • Arize/Phoenix
  • Weights & Biases tracing
  • custom dashboard in your app

These let you click a response and inspect:

  • query
  • retriever output
  • reranker output
  • prompt
  • final answer

Simple debugging workflow

  1. Ask a question.
  2. Look at the retrieved chunks.
  3. Verify the chunks contain the facts needed.
  4. Check whether the prompt includes those chunks.
  5. Inspect the citations in the answer.
  6. If the answer is wrong, determine whether the issue is:
    • retrieval failure
    • chunking issue
    • prompt issue
    • generation hallucination
    • citation formatting issue

If you want higher-quality citations

Use:

  • chunk-level citations
  • reranking
  • source quotation
  • answer grounding checks
  • “answer only from context” prompting
  • post-generation citation validation

Recommended minimum setup

At minimum, your app should save:

  • query
  • top_k retrieved chunks
  • source metadata
  • final answer
  • citation mapping

If you want, I can show you:

  1. a LangChain example,
  2. a LlamaIndex example, or
  3. a custom JSON logging format for citations and source chunks.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.