Prompt
Why did my RAG answers start citing the wrong documents after an index update?
Latest observation
This usually happens because the retrieval/index layer changed, but your citation mapping didn’t stay aligned. In a RAG system, “wrong documents” after an index update almost always comes from one of these causes:
Common reasons
1) Document IDs or chunk IDs got reassigned
If your index rebuild regenerated chunk IDs, the retriever may still return results in the new order, while your citation logic is looking up metadata using old IDs.
Symptom: answers cite documents that are real, but not the ones actually used.
2) Stale metadata cache
You may have updated the vector index but not the metadata store (or vice versa).
Symptom: retrieved chunk text is correct, citation title/source/date is outdated or mismatched.
3) Query-to-chunk mapping shifted due to re-embedding/rechunking
If you changed chunk size, overlap, embedding model, or normalization, nearest-neighbor results can change substantially.
Symptom: the system now retrieves different chunks, but the generator still produces citations from a previous pipeline assumption.
4) Index alias or pointer moved to a new version, but downstream code still reads the old one
For example, your search service may point to index_v2, while your citation resolver reads from index_v1.
Symptom: retrieval and citation resolution disagree on document sources.
5) Duplicate or near-duplicate chunks
After re-indexing, duplicate content may appear across documents or versions. The model may cite the wrong source because multiple chunks look equally relevant.
Symptom: citations are plausible but inconsistent across runs.
6) Mismatch between “retrieved chunk” and “cited source”
If you retrieve chunks, then summarize, then cite based on semantic similarity to the summary instead of the original retrieved chunks, the model can hallucinate citations.
Symptom: cited documents don’t actually contain the supporting text.
7) Filtering/sorting bug after update
A change in ranking, metadata filters, or tie-breaking can reorder top-k results. If your citation assignment assumes result order is stable, citations drift.
Symptom: the answer is okay, but footnote numbers or source labels point to the wrong item.
How to verify quickly
-
Log the exact retrieved chunk IDs and text
- For each answer, store the top-k retrieved chunks.
- Confirm the cited documents correspond to those exact chunks.
-
Check version consistency
- Ensure embeddings, vector index, metadata DB, and citation resolver all point to the same index version.
-
Compare old vs new chunking
- If chunk boundaries changed, old citation references may no longer map cleanly.
-
Inspect one bad example end-to-end
- Query → retrieved chunks → prompt passed to LLM → cited sources returned.
- This usually reveals where the mismatch begins.
-
Validate by source text
- For every citation, check whether the cited source contains the relevant span used in the answer.
Most likely root cause
If this started immediately after an index update, the most common issue is:
Your retriever and your citation metadata are out of sync due to changed document/chunk IDs or a stale index version reference.
Preventing it
- Use stable document IDs and stable chunk IDs.
- Store citations from the retrieved chunk metadata, not from regenerated guesses.
- Version your index and metadata together.
- Add an automated test that checks:
retrieved chunk → cited source → source text contains answer evidence.
If you want, I can help you diagnose it with a simple debugging checklist or by looking at your index/update pipeline design.