Prompt
I'm unhappy with my current vector DB because reindexing takes too long
Latest observation
Oct 7, 2026
OpenAI APIWeb search: off
If reindexing is the pain point, the best fix is often not just “a faster vector DB” but a system that supports incremental updates and low-cost rebuilds.
A few practical directions:
1) Use a vector store with fast upserts / incremental indexing
Look for:
- Near real-time upserts
- Background compaction / rebuilding
- Metadata + vector separation
- Sharding / partitioning
- Hybrid search support if you need keyword + vector
Examples people often evaluate:
- Qdrant
- Weaviate
- Pinecone
- Milvus
- OpenSearch / Elasticsearch kNN (if you already use them)
2) Reduce how often you need full reindexing
Common tricks:
- Store embeddings with stable document IDs and update only changed chunks
- Use content hashes to skip unchanged text
- Split large collections into logical partitions
- Keep old index live while building the new one, then swap alias/endpoint
- Re-embed only when the chunk text actually changes, not when parent docs change elsewhere
3) Consider architectural changes
If your data changes frequently:
- Use a hot index for recent updates and a cold index for the bulk corpus
- Periodically merge or compact in the background
- Use a two-stage retrieval pipeline:
- coarse ANN search
- rerank a smaller candidate set
4) If the issue is scale, look at operational knobs
Reindexing may be slow because of:
- too many tiny writes
- high-dimensional embeddings
- large payloads attached to vectors
- poor batching
- low HNSW/IVF tuning
- lack of parallelism
You can often improve things with:
- bigger batch sizes
- async ingestion
- tuned index parameters
- separate embedding generation from indexing
- fewer stored fields in the vector DB
5) Sometimes a different retrieval approach is better
If you don’t truly need ANN over a massive corpus, a simpler setup might be enough:
- Postgres + pgvector for moderate scale
- OpenSearch/Elasticsearch for hybrid retrieval
- FAISS if you want a library and can manage the ops yourself
If you want, I can help you narrow this down based on:
- your data size,
- how often documents change,
- current DB, and
- whether you need hybrid search or just vectors.