Prompt
I'm building a multilingual search feature and need help choosing an embedding strategy
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
A good multilingual embedding strategy usually depends on one question: do you want one shared semantic space for all languages, or language-specific spaces with translation/query routing?
Recommended default
For most search products, I’d start with one multilingual embedding model for both documents and queries.
Why this is usually best
- Cross-lingual retrieval: a Spanish query can find an English document if the model is truly multilingual.
- Operational simplicity: one index, one vector space, one pipeline.
- Better UX for mixed-language corpora: users can search across languages without needing language detection to reroute.
Good model characteristics
Look for models that:
- support many languages well,
- are trained for retrieval, not just similarity,
- handle short queries and long documents,
- have strong performance on your target languages.
Examples of model families to evaluate:
- bge-m3
- multilingual-e5
- LaBSE
- some commercial embedding APIs with multilingual support
When to use a different strategy
1) Language-specific embeddings per language
Use this if:
- your languages are very different in quality/coverage,
- you have heavy language-specific ranking needs,
- cross-lingual search is not important.
Pros:
- sometimes better monolingual relevance per language,
- can tune each language separately.
Cons:
- more complexity,
- harder cross-lingual search,
- more operational overhead.
2) Translation-based retrieval
Translate query into document language or all docs into a pivot language.
Use this if:
- you need very high relevance,
- your chosen embedding model is weak for some languages,
- you already have translation infrastructure.
Pros:
- can improve recall for low-resource languages.
Cons:
- latency, cost, translation errors,
- can distort named entities and domain terms.
3) Hybrid search: BM25 + embeddings
For production search, this is often the best overall approach.
Use:
- BM25 / keyword search for exact match, names, rare terms
- multilingual embeddings for semantic matching
- optional reranker for final ranking
This is especially strong for multilingual search because:
- embeddings help across languages,
- keyword search preserves exact-term precision,
- reranking improves relevance.
Practical recommendation
If I were designing this, I’d do:
- Use a multilingual retrieval embedding model
- Index all documents in one vector index
- Keep language metadata
- Combine vector search with keyword search
- Optionally rerank top results with a multilingual cross-encoder/reranker
- Evaluate separately for each language and cross-lingual pairs
How to choose the embedding model
Benchmark candidates on:
- same-language search
- cross-language search
- named entities
- domain-specific queries
- short vs long text
Measure:
- Recall@k
- nDCG@k
- MRR
- latency and cost
A simple decision rule
- Need cross-lingual search? → use a multilingual embedding model in a shared space.
- Need best precision and can afford complexity? → add hybrid retrieval + reranking.
- Need strong support for low-resource languages? → consider translation or language-specific tuning.
If you want, I can help you pick a strategy based on:
- your languages,
- corpus size,
- whether queries and docs are in the same language,
- latency/cost constraints,
- whether you want open-source or API-based embeddings.