Prompt
I'm building a RAG app and need to serve an embedding model, what are good options?
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
Good options depend mainly on your constraints: quality, latency, cost, language coverage, and whether you want to self-host or use an API. For a RAG app, these are common choices:
Strong hosted/API options
1) OpenAI text-embedding-3-small / text-embedding-3-large
- Very strong general-purpose embeddings
- Easy to use, reliable, fast
- Good default choice if API use is fine
3-smallis usually the cost-effective pick;3-largegives better retrieval quality
2) Cohere Embed
- Strong retrieval-oriented embeddings
- Good multilingual performance
- Often a solid choice for enterprise/search use cases
3) Voyage AI embeddings
- Known for strong retrieval quality
- Good option if you want top-tier embedding performance
Strong self-hosted/open models
4) BAAI bge family
- Popular open-source choice for RAG
- Good quality, widely used
- Examples:
bge-small,bge-base,bge-large, and newer variants likebge-m3for multilingual/multi-function use
5) e5 family
- Another widely used open embedding family
- Good for semantic search / retrieval
- Works well when fine-tuned or used with the recommended query/document prefixing patterns
6) Jina embeddings
- Good open/self-hostable options
- Often competitive for retrieval and multilingual tasks
7) SentenceTransformers models
- Huge ecosystem and lots of model choices
- Good if you want flexibility and easy local deployment
- Common starting points include models from
all-MiniLMfamily for lighter setups, though quality is lower than newer retrieval-focused models
What I’d recommend in practice
If you want the easiest high-quality setup:
- OpenAI
text-embedding-3-smallto start - Switch to
text-embedding-3-largeif retrieval quality matters more than cost
If you want to self-host:
bge-m3is a very good modern defaulte5-largeorbge-largeare also solid
How to choose
Consider:
- Quality: test on your own queries and docs
- Latency: smaller models are faster
- Cost: hosted APIs cost per token/vector; self-hosting costs infra
- Multilingual support:
bge-m3, Cohere, Voyage, and some multilingual models are strong - Domain fit: legal, code, support tickets, and technical docs may benefit from specific evaluation
Practical advice
For RAG, don’t just pick one blindly:
- Build a small eval set of real queries
- Measure retrieval metrics like recall@k / nDCG
- Test chunking + embedding model together
- If possible, compare:
text-embedding-3-smalltext-embedding-3-largebge-m3ore5-large
If you want, I can give you:
- a short list by budget
- a self-hosted deployment stack
- or a benchmark plan for choosing embeddings for your RAG app