Prompt
How do I keep embeddings updated when source documents change?
Latest observation
To keep embeddings in sync with changing source documents, you generally want an incremental re-embedding pipeline rather than reprocessing everything every time.
Common strategies
1. Detect changes in source docs
Track one or more of:
- Last modified timestamp
- Content hash of the document or chunk
- Version ID from the source system
If a document’s hash/version changes, mark it for reprocessing.
2. Chunk documents consistently
Split documents into chunks using a stable method so unchanged sections stay the same.
- Keep chunking rules deterministic
- Use stable chunk IDs if possible
- Store metadata like:
- document ID
- chunk index
- chunk hash
- embedding model version
This lets you update only the chunks that changed.
3. Re-embed only changed chunks
For each modified document:
- Recompute chunks
- Compare new chunk hashes to existing ones
- Re-embed only new or modified chunks
- Delete embeddings for removed chunks
4. Version your embeddings
Store the model and pipeline version with each embedding:
embedding_model = text-embedding-3-largemodel_version = 2026-01chunking_version = v2
If you change the embedding model or chunking strategy, you may need a full reindex.
5. Use upserts in your vector database
Most vector stores support:
- upsert for new/updated vectors
- delete for stale vectors
A typical flow:
- Fetch changed document
- Re-chunk
- Embed new/changed chunks
- Upsert updated vectors
- Delete chunks no longer present
6. Maintain an index manifest
Keep a table or metadata store with:
- document ID
- chunk ID
- content hash
- embedding status
- last indexed time
- source version
This helps you audit and sync reliably.
Recommended architecture
A practical setup looks like this:
-
Source change detector
- polls files, listens to webhooks, or watches a database
-
Chunking service
- turns docs into stable chunks
-
Embedding queue
- processes only changed chunks asynchronously
-
Vector store
- upserts embeddings and removes obsolete ones
-
Metadata store
- tracks doc/chunk hashes and versions
If you need a simple rule
A good default is:
- If the document hash changed, re-chunk it
- If a chunk hash changed, re-embed it
- If a chunk disappeared, delete its vector
- If the embedding model changed, re-embed everything
Example workflow
Suppose you have a doc with 10 chunks. After editing:
- Chunks 1–7 unchanged
- Chunk 8 edited
- Chunk 9 deleted
- Chunk 10 unchanged
You would:
- keep vectors for 1–7 and 10
- re-embed chunk 8
- delete chunk 9
- upsert the new chunk 8 vector
Extra tips
- Use idempotent jobs so retries don’t duplicate vectors
- Batch embeddings for efficiency
- Separate source-of-truth text from derived embeddings
- Keep raw chunk text stored so you can re-embed later without re-fetching the source
If you want, I can also show:
- a database schema for tracking embeddings, or
- a Python example for incremental reindexing.