Prompt
I’m building a cross-lingual search feature and need the right embedding setup
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
For a cross-lingual search feature, the best setup is usually:
Recommended default
Use a multilingual bi-encoder embedding model that maps text from different languages into the same vector space.
That lets you:
- index documents in one language
- search with queries in another
- retrieve semantically similar content across languages
What to look for in the model
Choose an embedding model that is:
- multilingual: supports the languages you need
- semantic: works well for paraphrase / meaning matching, not just translation
- retrieval-optimized: trained for search / relevance
- consistent dimensionality across languages
- ideally supports longer text if your documents are not short
Good model categories
1) Multilingual sentence embeddings
Best for general cross-lingual retrieval.
Examples:
- LaBSE
- multilingual-e5
- bge-m3
- distiluse-base-multilingual-cased-v2
2) Retrieval-tuned multilingual embeddings
Often better for search than generic sentence embeddings.
Examples:
- multilingual-e5-large
- bge-m3
These usually work well with query/document formatting like:
query: ...passage: ...
Suggested architecture
For document search:
- Embed all documents offline
- Store vectors in a vector index
- Embed the user query at runtime
- Retrieve top-K nearest neighbors
- Optionally rerank with a cross-encoder or LLM reranker
Practical advice
- If your search is mostly multilingual semantic retrieval, start with bge-m3 or multilingual-e5.
- If you need many languages and strong zero-shot behavior, LaBSE is a solid baseline.
- If you also care about lexical matching, consider a hybrid search setup:
- dense embeddings + BM25
- then rerank results
Important tuning points
- Use cosine similarity for most embedding search setups
- Normalize embeddings if the model expects it
- Test with your real query/document pairs; multilingual quality varies a lot by language pair
- Evaluate separately for:
- same-language retrieval
- cross-language retrieval
- low-resource languages
If you want the simplest strong default
I’d suggest:
- Model:
bge-m3ormultilingual-e5-large - Index: vector DB with cosine similarity
- Retrieval: top 20–100
- Reranking: optional cross-encoder if latency allows
- Search mode: hybrid dense + sparse if precision matters
If you want, I can also help you choose between:
- open-source vs hosted embeddings,
- the best model for your language set, or
- a concrete indexing/query pipeline for your stack.