Prompt
semantic search latency too high embeddings pipeline
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
If your semantic search latency is too high, the bottleneck is usually one of these parts of the embeddings pipeline:
- Embedding generation at query time
- Vector database / ANN retrieval
- Post-filtering / reranking
- Network / serialization overhead
- Poor index configuration or missing caching
Fast ways to reduce latency
1) Precompute everything you can
- Documents: embed offline, not during the query path.
- Queries: if repeated or templated, cache query embeddings.
- Store embeddings in a vector index ahead of time.
2) Use a smaller/faster embedding model
- Choose a model optimized for retrieval, not general generation.
- Smaller dimensionality often helps both embedding time and search time.
- Consider batching for offline ingestion, but keep query-time embedding lightweight.
3) Add a cache layer
- Cache:
- query embeddings
- top-k results for frequent queries
- reranker outputs for common queries
- Even a short TTL cache can cut p95 latency a lot.
4) Tune ANN index settings
If using FAISS / HNSW / Milvus / Pinecone / Weaviate / Qdrant:
- Lower
efSearch/ probe counts if recall is acceptable - Ensure the index type matches your scale:
- HNSW: good low-latency search
- IVF / PQ: good for very large collections, but needs tuning
- Make sure vectors are normalized if using cosine similarity.
5) Reduce candidate set before reranking
- Retrieve top 20–100 from vector search
- Only rerank a small subset
- Avoid reranking hundreds/thousands unless necessary
6) Keep filtering cheap
- Apply metadata filters in the vector DB if supported
- Avoid expensive application-side filtering after retrieval
- Pre-partition data by tenant / language / category if applicable
7) Optimize ingestion and storage
- Use float16 / quantized vectors if accuracy allows
- Batch writes
- Avoid reconstructing vectors on read
- Keep hot indexes in memory
8) Measure the right latency segments
Break latency into:
- embedding generation
- vector search
- metadata filter
- reranking
- response serialization
This helps identify whether the fix is in the model, index, or app.
Typical low-latency architecture
Query → embed → vector search → optional rerank → return
To make this fast:
- query embedding model on GPU or optimized CPU runtime
- ANN index in memory
- rerank only top few
- cache heavily
Common anti-patterns
- Computing embeddings for documents at query time
- Using a large embedding model for every request
- Searching a very large unsharded index with weak ANN tuning
- Reranking too many candidates
- Filtering after retrieval instead of inside the index
- No caching for repeat queries
If you want, I can help you with one of these:
- diagnose your current pipeline bottleneck
- suggest an optimized architecture
- tune FAISS / HNSW / Pinecone / Milvus / Weaviate / Qdrant
- estimate expected latency from your scale
If you paste your stack and metrics, I can be specific. For example:
- embedding model
- vector DB / index type
- corpus size
- p50/p95 latency
- top-k
- reranking yes/no