Prompt
I'm building a retrieval app and need low-latency model serving
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
For a retrieval app, low-latency serving usually means optimizing around embedding generation, reranking, and sometimes a small LLM for answer synthesis.
Practical low-latency setup
-
Use a small/fast embedding model
- Good default:
bge-small,e5-small, or similar compact sentence-transformer style models. - If quality matters more, move up to
bge-base/e5-base, but latency increases.
- Good default:
-
Batch requests
- If traffic is bursty, micro-batching can improve throughput without much latency penalty.
- Keep batch sizes small for interactive workloads.
-
Serve on GPU if volume is moderate/high
- Use TensorRT, ONNX Runtime, or optimized inference stacks if available.
- For CPU-only, quantized models can help a lot.
-
Use ANN retrieval
- FAISS, HNSW, ScaNN, or a managed vector DB with HNSW/IVF indexes.
- Precompute embeddings for documents; only query-time embeddings should be generated online.
-
Separate retrieval stages
- Stage 1: fast vector search
- Stage 2: optional reranker on top-k only
- Stage 3: optional LLM response generation
- This keeps most queries fast and only spends more compute on hard cases.
Latency tips
- Keep payloads small.
- Avoid heavy prompt construction at query time.
- Warm up models on startup.
- Pin model replicas close to the vector DB region.
- Use async I/O and caching for repeated queries.
Serving stacks to consider
- vLLM: great for LLM serving, especially if you need generation.
- TGI (Text Generation Inference): solid for transformer serving.
- ONNX Runtime / OpenVINO: strong for low-latency CPU inference.
- Triton Inference Server: flexible for multi-model deployment.
- SentenceTransformers + ONNX/TorchScript: common for embeddings.
Recommended architecture
- Embedding service: small transformer, GPU or optimized CPU
- Vector index: FAISS/HNSW
- Reranker service: only for top 20–50 candidates
- Generator: optional, only if you need natural-language answers
If you want, I can suggest:
- a specific model + stack for your hardware, or
- a reference architecture with latency targets and component choices.