Prompt
I’m unhappy with Elasticsearch semantic search quality
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
Sorry you’re dealing with that — semantic search quality in Elasticsearch can be good, but it’s easy for it to underperform if the embeddings, indexing strategy, or query flow aren’t tuned.
A few common causes:
- Embedding model mismatch: the model may not fit your domain or query style.
- Chunking issues: documents may be split too large or too small.
- No hybrid search: semantic-only often misses exact keyword matches.
- Wrong similarity / field setup: vector field mapping, normalization, or kNN parameters may be off.
- Poor query formulation: the query text may not match how the embeddings were trained.
- Insufficient re-ranking: top results from kNN may need a second-stage ranker.
Typical ways to improve:
-
Use hybrid retrieval
- Combine BM25 keyword search + vector search.
- This often gives a big boost over semantic-only.
-
Improve chunking
- Aim for semantically coherent chunks, not arbitrary token splits.
- Keep chunks small enough that each one has a clear topic.
-
Try a better embedding model
- Especially if your content is technical, legal, product-specific, or multilingual.
-
Add re-ranking
- Use a cross-encoder or LLM-based reranker on the top 20–100 candidates.
-
Tune kNN settings
- Increase candidate pool, adjust
num_candidates, and test recall vs latency.
- Increase candidate pool, adjust
-
Evaluate with a small labeled set
- Create 20–100 representative queries and judge top results.
- This helps identify whether the problem is retrieval, chunking, or ranking.
If you want, I can help you diagnose it systematically. For example, I can walk you through:
- a debug checklist
- a hybrid search setup
- better chunking strategies
- or how to evaluate Elasticsearch semantic search quality
If you share:
- your Elasticsearch version,
- the embedding model,
- your mapping,
- and an example query + bad results,
I can suggest concrete fixes.
Brands