Prompt
What's the best architecture for storing embeddings, updating them often, and keeping retrieval latency low?
Latest observation
The “best” architecture usually depends on scale and update rate, but for frequent embedding updates + low retrieval latency, the most reliable pattern is:
Recommended architecture: Hybrid online vector index + immutable snapshot + async rebuild/merge
Use:
-
Primary vector store for serving reads
- An ANN index like HNSW, IVF-PQ, or a managed vector DB.
- Optimized for low-latency similarity search.
-
Write path decoupled via a queue/stream
- New or updated embeddings go into a message queue (Kafka, Pulsar, SQS, etc.).
- A background indexing service consumes updates and applies them asynchronously.
-
Immutable segment storage / snapshotting
- Keep vectors in append-only segments or periodic snapshots.
- Rebuild/compact indices in the background, then atomically swap them in.
-
Hot update overlay
- For very fresh embeddings, maintain a small delta index in memory.
- Query both the main index and delta index, then merge results.
- Periodically fold the delta into the main index.
-
Metadata store separate from vector store
- Keep IDs, timestamps, ACLs, filters, and document metadata in a transactional DB (Postgres, MySQL, etc.).
- Vector DB stores vector + pointer/reference, not all business state.
Why this works
Low latency
ANN indexes like HNSW are excellent for fast approximate nearest-neighbor retrieval.
Frequent updates
Many vector indexes are expensive to mutate in place. An append + background compaction approach avoids blocking queries.
Consistency without slowing reads
Writes are eventually consistent, while reads stay fast and stable.
Best choices by workload
1) If updates are frequent and latency matters most
HNSW + delta overlay
- Best for online systems with lots of inserts/updates/deletes.
- Query latency is usually very good.
- Downside: memory-heavy at large scale.
2) If corpus is huge and mostly read-heavy
IVF-PQ / DiskANN / managed vector DB with background indexing
- Better storage efficiency.
- Lower memory cost.
- Usually a bit more operational complexity and often slightly higher latency than HNSW.
3) If vectors change constantly and freshness is critical
Two-tier retrieval
- Tier 1: small in-memory “fresh” index
- Tier 2: larger durable ANN index
- Merge results at query time
This is often the sweet spot.
Practical architecture pattern
Write path
- Application updates embedding
- Store raw object + metadata in primary DB
- Publish embedding update event
- Indexer updates:
- delta index immediately
- main index asynchronously
Read path
- Query delta index
- Query main ANN index
- Merge top-k results
- Apply metadata filtering / reranking
Important design tips
- Use versioned embeddings so you can replace stale ones safely.
- Prefer soft deletes + background cleanup over immediate physical deletion.
- Batch updates where possible to reduce index churn.
- If you filter heavily by metadata, consider:
- pre-filtered partitions, or
- hybrid search with metadata-aware routing.
- If exact freshness matters, store the latest embedding in a fast key-value store and reconcile asynchronously.
Rule of thumb
If you want a broadly good answer:
- HNSW for low latency
- append-only updates
- small hot delta index
- async background compaction/rebuild
- separate metadata DB
That’s usually the best balance of freshness, latency, and operational simplicity.
If you want, I can also give you:
- a reference architecture diagram,
- specific tech stacks for small/medium/large scale, or
- a comparison of HNSW vs IVF vs DiskANN vs FAISS for this use case.