Prompt
What should I use for embeddings if I need low latency at query time?
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
If low query-time latency is the priority, use small, fast embedding models rather than the largest “best quality” ones.
Good general choices:
- OpenAI
text-embedding-3-small— a strong default for speed/cost/quality - Sentence-Transformers MiniLM models — e.g.
all-MiniLM-L6-v2 - E5-small / E5-base — good retrieval embeddings with decent speed
- BGE-small — another solid low-latency option
Practical guidance:
- If you want fastest + good enough: start with
text-embedding-3-smallor MiniLM - If you need higher recall but still reasonable latency: try E5-base or BGE-base
- If latency is critical, also consider:
- precomputing embeddings
- using a vector index like HNSW/FAISS/ScaNN
- smaller dimensions if your model supports them
- batching requests where possible
Rule of thumb:
- Embedding generation usually dominates latency if done on the fly.
- For query-time speed, the biggest wins are often smaller models + caching + ANN search.
If you want, I can recommend a specific embedding model based on your stack (OpenAI, local GPU/CPU, Python, vector DB, etc.).