Prompt
I'm building a retrieval-augmented app and want to see which chunks were used
Latest observation
Oct 9, 2026
OpenAI APIWeb search: off
To see which chunks were used in a retrieval-augmented generation (RAG) app, you typically need to surface the retrieval step’s results alongside the final answer.
What to track
For each user query, log:
- Retrieved chunk text
- Chunk metadata: source doc, page/section, chunk id, score
- Prompt sent to the model
- Final answer
- Optional: re-ranking results and citations
Common ways to expose it
-
Return retrieved chunks in the API response
- Useful for debugging and transparency.
- Example response shape:
{ "answer": "…", "retrieved_chunks": [ { "chunk_id": "doc1_03", "source": "doc1.pdf", "score": 0.82, "text": "..." } ] } -
Add citations in the answer
- e.g.
[1],[2]linked to chunk metadata.
- e.g.
-
Store retrieval traces
- Save query, top-k results, and prompt context in a database or log system.
-
Use a tracing/observability tool
- LangSmith, Weights & Biases, OpenTelemetry, etc.
Minimal implementation pattern
If you’re using a vector DB, your flow is usually:
results = vector_store.similarity_search(query, k=5)
context = "\n\n".join([r.page_content for r in results])
answer = llm.generate(prompt=query + context)
return {
"answer": answer,
"retrieved_chunks": [
{
"text": r.page_content,
"metadata": r.metadata
}
for r in results
]
}
If you want to inspect them during development
- Print the top-k results before generation
- Log similarity scores
- Save them to a file or dashboard
- Show a “sources” panel in your UI
If you tell me what stack you’re using — e.g. LangChain, LlamaIndex, OpenAI Assistants, Pinecone, FAISS, Chroma, Weaviate — I can show exactly how to extract and display the chunks.