Prompt
what embedding model should I use for multilingual search?
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
For multilingual search, the best embedding model depends on your priorities:
Good default choice
intfloat/multilingual-e5-large
- Strong multilingual retrieval performance
- Works well for semantic search across many languages
- Commonly used and easy to adopt
- Best when you want one solid model for query/document search
Other strong options
BAAI/bge-m3
- Very strong multilingual model
- Supports dense, sparse, and multi-vector retrieval
- Good if you want flexibility or better retrieval quality in more complex setups
intfloat/multilingual-e5-base
- Smaller/faster than
large - Good if you need lower latency or lower cost
- Slightly less accurate than the large version
If you need hosted/API models
Look at:
- OpenAI text-embedding-3-large or text-embedding-3-small
- These are strong and multilingual, but if you need local/open-source deployment, the models above are better fits.
Practical recommendation
- Best overall open-source default:
intfloat/multilingual-e5-large - Best if you want state-of-the-art flexibility:
BAAI/bge-m3 - Best lightweight option:
intfloat/multilingual-e5-base
Important tip
For multilingual search, results often improve a lot if you:
- use the same embedding model for both queries and documents
- normalize text consistently
- test with your actual languages and search data
- consider reranking after retrieval for better relevance
If you want, I can recommend a model based on:
- your languages,
- whether you need open-source vs API, and
- your latency/cost constraints.