Prompt
I'm building a system that updates embeddings every few minutes, what should I use?
Latest observation
Oct 7, 2026
OpenAI APIWeb search: off
If you need to update embeddings every few minutes, you’ll usually want a setup that supports frequent re-embedding, fast upserts, and incremental indexing rather than a batch-only pipeline.
Good options
1. Vector database with upsert support
Best if you need embeddings searchable in near real time.
Examples:
- Pinecone
- Weaviate
- Qdrant
- Milvus
- pgvector (if scale is moderate and you already use Postgres)
Why:
- Supports frequent updates
- Handles ANN indexing
- Simplifies retrieval pipeline
2. Event-driven embedding pipeline
If the data changes often, use a queue + worker pattern:
- Kafka / RabbitMQ / SQS
- Embedding worker regenerates vectors
- Upserts into vector store
Good when:
- You want updates every few minutes
- Data arrives continuously
- You need retries and backpressure handling
3. Incremental re-embedding strategy
Don’t recompute everything.
- Only re-embed changed/new items
- Track content hashes or timestamps
- Use background jobs for stale items
This is usually the biggest performance win.
If you’re choosing a stack
- Small/medium scale:
Postgres + pgvector + background job queue - Production scale / high QPS:
QdrantorPinecone - Already on cloud infra:
SQS + Lambda/ECS workers + vector DB
Practical recommendation
If embeddings update every few minutes, I’d recommend:
Queue-based incremental embedding pipeline + vector database with upserts
A common architecture:
- Source data changes
- Change event goes to queue
- Worker generates embedding
- Upsert into vector DB
- Retrieval service queries vector DB
Things to watch
- Index refresh latency
- Consistency between source data and vectors
- Deduping repeated updates
- Batching embeddings for efficiency
- Versioning embeddings if model changes
If you want, I can recommend a specific stack based on:
- your scale,
- cloud provider,
- whether updates are by record or by document,
- and whether you need real-time search or just periodic refresh.