Prompt

I’m unhappy with Elasticsearch semantic search quality

Artificial Intelligence / AI Search1 observationLast seen Oct 10, 2026

Latest observation

Oct 10, 2026OpenAI APIWeb search: off

Sorry you’re dealing with that — semantic search quality in Elasticsearch can be good, but it’s easy for it to underperform if the embeddings, indexing strategy, or query flow aren’t tuned.

A few common causes:

  • Embedding model mismatch: the model may not fit your domain or query style.
  • Chunking issues: documents may be split too large or too small.
  • No hybrid search: semantic-only often misses exact keyword matches.
  • Wrong similarity / field setup: vector field mapping, normalization, or kNN parameters may be off.
  • Poor query formulation: the query text may not match how the embeddings were trained.
  • Insufficient re-ranking: top results from kNN may need a second-stage ranker.

Typical ways to improve:

  1. Use hybrid retrieval

    • Combine BM25 keyword search + vector search.
    • This often gives a big boost over semantic-only.
  2. Improve chunking

    • Aim for semantically coherent chunks, not arbitrary token splits.
    • Keep chunks small enough that each one has a clear topic.
  3. Try a better embedding model

    • Especially if your content is technical, legal, product-specific, or multilingual.
  4. Add re-ranking

    • Use a cross-encoder or LLM-based reranker on the top 20–100 candidates.
  5. Tune kNN settings

    • Increase candidate pool, adjust num_candidates, and test recall vs latency.
  6. Evaluate with a small labeled set

    • Create 20–100 representative queries and judge top results.
    • This helps identify whether the problem is retrieval, chunking, or ranking.

If you want, I can help you diagnose it systematically. For example, I can walk you through:

  • a debug checklist
  • a hybrid search setup
  • better chunking strategies
  • or how to evaluate Elasticsearch semantic search quality

If you share:

  • your Elasticsearch version,
  • the embedding model,
  • your mapping,
  • and an example query + bad results,

I can suggest concrete fixes.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.