Prompt
What’s the best setup for multilingual embeddings and semantic search across markets?
Latest observation
The best setup is usually one shared multilingual embedding model + language-aware indexing/retrieval + reranking + market-specific tuning. That gives you good cross-lingual semantic search without maintaining separate systems per market.
Recommended architecture
1) Use a strong multilingual embedding model
Pick a model that embeds many languages into the same vector space so queries in one language can retrieve documents in another.
Good options:
- OpenAI text-embedding-3-large or -3-small for broad multilingual coverage
- Cohere Embed multilingual
- LaBSE / mE5 / multilingual-e5 for open-source stacks
- bge-m3 if you want a strong open-source retrieval model with multilingual support
Rule of thumb:
- If you want simplest, highest-quality managed setup: use a hosted multilingual embedding API.
- If you need on-prem / self-hosted: use bge-m3 or multilingual-e5.
2) Normalize content before embedding
For all markets, standardize:
- HTML cleanup
- language detection
- OCR cleanup if needed
- deduplication
- consistent chunking
- metadata tagging by:
- language
- market/region
- product/category
- source
- timestamp
This improves retrieval more than people expect.
3) Chunk by meaning, not just length
For documents, use semantic chunking:
- 200–500 tokens per chunk is a good starting point
- preserve headings, titles, and nearby context
- store parent-child relationships so you can return a relevant section and the full doc if needed
If your market content is short-form (FAQs, product cards, reviews), you may index the full item instead of chunking.
4) Store vectors with metadata filters
Use a vector database or search engine that supports:
- ANN vector search
- metadata filtering
- hybrid search
- reranking integration
Good choices:
- Pinecone
- Weaviate
- Milvus
- pgvector for simpler systems
- Elasticsearch/OpenSearch if you want hybrid lexical + semantic in one place
For cross-market setups, metadata filters matter a lot:
market = DElanguage in [de, en]category = electronics
This prevents noisy retrieval from irrelevant markets while still allowing cross-lingual matches when useful.
5) Use hybrid retrieval
Best practice is usually:
- BM25 / keyword search
- vector search
- merge results
- then rerank
Why:
- embeddings are great for meaning
- keyword search is great for exact terms, SKUs, legal wording, and brand names
- combining both improves recall
This is especially important across markets because:
- names, acronyms, and product codes often don’t translate well
- local terminology may differ even for the same concept
6) Add a reranker
A cross-encoder reranker improves precision dramatically.
Workflow:
- retrieve top 50–200 candidates
- rerank top results with a multilingual reranker
- return top 5–10
Options:
- managed rerankers from Cohere or similar
- open-source multilingual rerankers like bge-reranker variants
This is often the difference between “pretty good” and “production-grade.”
7) Consider language-aware query handling
For multilingual search across markets:
- detect query language
- optionally translate the query into a canonical language
- or search in native language + embedding space
Best practice:
do not rely only on translation. Use multilingual embeddings directly, and translation as a fallback/enhancement.
Useful pattern:
- query in original language
- vector search across multilingual corpus
- if query is short or ambiguous, expand with translated or synonym variants
- rerank using the original query language
8) Decide between one global index vs regional indexes
This depends on how different your markets are.
One global index
Best when:
- products/content overlap across markets
- users may search across languages
- you want unified analytics and fewer systems
Regional indexes
Best when:
- regulations or catalogs differ significantly
- latency needs are strict
- you have large market-specific content volumes
- content should not cross certain boundaries
Common compromise
- maintain one global semantic index
- use metadata filters to constrain by market when needed
- optionally maintain a few regional shards for performance/governance
Best-practice stack by maturity
Simple, high-quality production setup
- Multilingual embeddings: managed API
- Vector DB: Pinecone / Weaviate / OpenSearch
- Hybrid retrieval: yes
- Reranking: yes
- Metadata filters: yes
Open-source stack
- Embeddings: bge-m3 or multilingual-e5
- Vector DB: Milvus / Weaviate / pgvector
- Lexical search: OpenSearch
- Reranker: multilingual bge-reranker
- Pipeline: Python + batch indexing jobs
Enterprise/global commerce setup
- One canonical content pipeline
- language detection + normalization
- global vector index + regional filters/shards
- hybrid retrieval
- multilingual reranking
- query logging + human evals per market
- feedback loop for market-specific tuning
Important design choices
Embedding model selection
Evaluate with your actual data:
- short queries
- product discovery
- support FAQs
- legal/policy search
- brand/product name retrieval
A model that scores well on benchmarks may still fail on:
- local abbreviations
- transliterated names
- code-mixed queries
- domain-specific jargon
Cross-market retrieval strategy
If the same concept has different wording across markets:
- add synonyms/aliases per locale
- store translations of titles and key fields
- enrich metadata with canonical product/category IDs
- keep locale-specific text separate from canonical entities
Example:
- canonical concept: “wireless earbuds”
- market terms:
- US: “wireless earbuds”
- DE: “Bluetooth-Kopfhörer”
- JP: “ワイヤレスイヤホン”
Use canonical IDs plus multilingual text so the system can connect them.
Evaluation
Measure by market and language, not just globally:
- Recall@K
- MRR
- nDCG
- query success rate
- zero-result rate
- click-through / conversion
Build a test set with:
- native-language queries
- cross-language queries
- brand/SKU queries
- slang and regional terms
- ambiguous queries
A practical default recommendation
If you want a strong starting point:
- Embedding model: multilingual embedding model with strong cross-lingual support
- Index: single global vector index with market/language metadata
- Retrieval: hybrid search
- Rerank: multilingual reranker
- Chunking: semantic, 200–500 tokens
- Filters: market, language, category, recency
- Evaluation: per-market offline test set + online logging
If you want the shortest answer
The best setup is usually:
Multilingual embeddings + hybrid search + reranking + metadata filters + per-market evaluation, with either:
- one global index if content overlaps heavily, or
- regional shards if markets are very different.
If you want, I can also recommend a specific stack based on your constraints:
- hosted vs open-source,
- budget,
- number of languages/markets, and
- whether you’re searching products, docs, support tickets, or legal content.