Prompt
What should I use for embeddings if my team wants to self-host?
Latest observation
Oct 7, 2026
OpenAI APIWeb search: off
If your team wants to self-host embeddings, the best choice usually depends on your constraints, but the safest defaults are:
Good self-hosted options
1. bge-small / bge-base / bge-large
From BAAI. These are very popular for self-hosted semantic search and RAG.
- Pros: strong retrieval quality, widely used, easy to deploy
- Cons: slightly more setup if you want best performance
- Best for: general-purpose document search and RAG
Common picks:
bge-small-en-v1.5— fast, cheap, good baselinebge-base-en-v1.5— strong balancebge-large-en-v1.5— better quality, more compute
2. E5 models
Also very solid for retrieval.
- Pros: high quality, especially for search tasks
- Cons: requires using the right query/document formatting
- Best for: search, retrieval, multilingual use cases
Examples:
e5-small-v2e5-base-v2multilingual-e5-base
3. Jina embeddings
Good if you want a modern option with strong multilingual support.
- Pros: good quality, convenient
- Cons: less “standard” than BGE/E5 in some teams
- Best for: multilingual or mixed workloads
4. Sentence Transformers models
If you want lots of flexibility and easy self-hosting, this ecosystem is very practical.
- Pros: huge model catalog, simple deployment
- Cons: quality varies by model
- Best for: prototyping and custom workloads
If you want a single recommendation
- English-only general RAG:
bge-base-en-v1.5 - Need faster/cheaper:
bge-small-en-v1.5 - Multilingual:
multilingual-e5-baseor a multilingual BGE model - Highest quality with more compute:
bge-large-en-v1.5
Self-hosting stack suggestions
You can serve embeddings with:
- SentenceTransformers + FastAPI
- vLLM if using supported embedding models
- Hugging Face TEI (Text Embeddings Inference) for a production-ready server
- ONNX / TensorRT if you want lower latency
What to optimize for
When choosing, ask:
- Language coverage — English only or multilingual?
- Latency — do you need fast CPU inference?
- Accuracy — retrieval quality vs cost
- Context size / input length — long docs may matter
- GPU availability — big impact on model choice
Practical default
If you’re unsure, start with:
bge-base-en-v1.5 via Hugging Face TEI
It’s a strong, self-hostable baseline that works well for most RAG/search use cases.
If you want, I can also give you:
- a recommended model by use case
- a CPU-only self-hosting setup
- or a Docker example for embedding inference