Prompt
re-embedding pipeline for updated documents
Latest observation
Here’s a practical re-embedding pipeline for updated documents. This assumes you already have a vector store / embeddings index and want to keep it in sync when documents change.
Goals
- Detect document changes
- Re-chunk updated content
- Re-embed only what changed when possible
- Replace stale vectors safely
- Keep metadata and versioning consistent
Recommended pipeline
1. Ingest documents
Pull documents from your source:
- files, S3, DB, CMS, Git repo, web pages, etc.
Store:
doc_idsource_uricontentlast_modifiedcontent_hashversion
2. Detect updates
Compare incoming docs against indexed docs using:
last_modifiedcontent_hash- document version number
- diff against previous content
If unchanged, skip.
If changed:
- mark previous vectors as stale
- enqueue for reprocessing
3. Re-process document text
Run the same preprocessing as initial ingestion:
- normalize text
- remove boilerplate if needed
- extract structured sections
- preserve headers and hierarchy if useful
Important: keep preprocessing deterministic so embeddings stay comparable.
4. Re-chunk document
Recreate chunks from the updated content.
Best practice:
- chunk by semantic boundaries when possible
- keep chunk size stable
- include overlap if helpful
- attach metadata:
doc_idchunk_idchunk_indexsection_pathcontent_hashversion
If chunk boundaries changed significantly, treat the whole doc as re-chunked, not just edited chunks.
5. Embed chunks
Generate embeddings for all new chunks.
Optimization options:
- only embed chunks whose hash changed
- parallelize embedding requests
- batch embeddings for throughput
6. Upsert into vector store
For each chunk:
- insert new vector record
- update metadata
- ensure idempotency with stable IDs
A good pattern is:
- vector ID =
{doc_id}:{chunk_hash}or{doc_id}:{version}:{chunk_index}
If using hash-based IDs, unchanged chunks can be reused.
7. Delete stale vectors
Remove old embeddings for the same doc/version range.
Options:
- hard delete old vectors after successful upsert
- soft delete via
active=false - keep prior versions for audit/history
If you support document rollback, soft delete or versioned retention is useful.
8. Rebuild derived indexes if needed
If you have:
- hybrid search indexes
- BM25 / keyword indexes
- reranker caches
- summary stores
rebuild or refresh them too.
Minimal architecture
Source docs
↓
Change detector
↓
Preprocess + chunk
↓
Embed
↓
Upsert vector store
↓
Delete stale vectors
↓
Refresh search/index metadata
Good design choices
Use content hashes
Hash each chunk’s normalized text:
- avoids re-embedding unchanged chunks
- helps deduplication
Version your documents
Keep doc_version or source_revision so you can:
- audit changes
- rollback
- trace query results to a specific revision
Make the pipeline idempotent
If the job runs twice, it should not duplicate embeddings.
Use transactional semantics if possible
Especially:
- write new vectors first
- validate success
- then delete old vectors
Example pseudo-code
def reembed_document(doc):
old_doc = metadata_store.get(doc.id)
new_hash = hash_text(doc.content)
if old_doc and old_doc.content_hash == new_hash:
return "unchanged"
chunks = chunk_text(preprocess(doc.content))
new_records = []
for i, chunk in enumerate(chunks):
chunk_hash = hash_text(chunk)
record_id = f"{doc.id}:{chunk_hash}"
if not vector_store.exists(record_id):
embedding = embed(chunk)
new_records.append({
"id": record_id,
"embedding": embedding,
"metadata": {
"doc_id": doc.id,
"chunk_index": i,
"chunk_hash": chunk_hash,
"version": doc.version
}
})
vector_store.upsert(new_records)
if old_doc:
vector_store.delete_where(doc_id=doc.id, version=old_doc.version)
metadata_store.save({
"doc_id": doc.id,
"content_hash": new_hash,
"version": doc.version,
"updated_at": now()
})
return "reembedded"
Practical improvements
- Use a queue like SQS/Kafka/RabbitMQ for update events
- Add retry logic for embedding API failures
- Store embedding model name/version in metadata
- Re-embed all docs if you switch embedding models
- Add monitoring for:
- docs processed
- chunks embedded
- failures
- latency
- stale vector deletions
If you need this for a production system
A robust production approach is:
- Detect changed doc
- Create new doc revision
- Chunk + embed revision
- Upsert new vectors
- Atomically switch active revision
- Garbage collect old revision later
This avoids search downtime and partial updates.
If you want, I can also give you:
- a SQL schema for tracking document versions and chunk hashes,
- a Python implementation, or
- a LangChain / LlamaIndex re-embedding workflow.