Prompt
What should I use for a cheap embeddings pipeline with acceptable recall?
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
If you want cheap embeddings with “good enough” recall, the best choice is usually:
Recommended default
sentence-transformers/all-MiniLM-L6-v2
- Very cheap / fast
- Small model (~22M params)
- Works well for general semantic search
- Common baseline for production prototypes and budget systems
Why it’s a good fit
- High throughput on CPU
- Low memory use
- Strong quality for its size
- Easy to deploy locally or on a modest API budget
If you want a bit better recall
BAAI/bge-small-en-v1.5
- Often better retrieval quality than MiniLM
- Still relatively cheap
- Good choice for English-only search
If you can afford slightly more compute, this is often the better “cheap but decent” option.
If multilingual is needed
intfloat/multilingual-e5-small
- Good for multilingual retrieval
- Still lightweight
- Use the right prompting format for E5 models, e.g.:
- query:
"query: ..." - document:
"passage: ..."
- query:
Practical pipeline for cheap + acceptable recall
If cost matters, don’t overspend on embeddings alone. Use:
- Small embedding model
- MiniLM or BGE-small
- ANN vector index
FAISS,hnswlib, orpgvector
- Chunking
- Use sensible chunk sizes, e.g. 200–500 tokens
- Reranking for top results
- Cheap embeddings + a small cross-encoder reranker on top 20–50 results can improve quality a lot
This usually gives much better recall/precision tradeoff than using a large embedding model everywhere.
If you want the absolute cheapest acceptable setup
- Model:
all-MiniLM-L6-v2 - Index: FAISS HNSW or IVF
- Reranker: optional, only if quality matters
- Chunk size: 300–400 tokens with overlap
Quick rule of thumb
- Lowest cost:
all-MiniLM-L6-v2 - Best cheap quality:
bge-small-en-v1.5 - Multilingual cheap:
multilingual-e5-small
If you tell me:
- your language(s),
- data size,
- latency budget,
- and whether this is for semantic search / RAG / recommendations,
I can suggest a more specific stack.