Prompt

OpenAI embeddings are not giving me good retrieval results

Artificial Intelligence / AI Search1 observationLast seen Oct 10, 2026

Latest observation

Oct 10, 2026OpenAI APIWeb search: off

That usually comes down to one of a few issues: chunking, query formulation, index settings, or expectation mismatch—not necessarily the embedding model itself.

Here’s a practical checklist to improve retrieval:

1) Fix chunking first

Bad chunking is the most common cause.

  • Too large: chunks contain multiple topics, so embeddings get “blurred.”
  • Too small: chunks lose context and become ambiguous.
  • Good starting point:
    • 200–500 tokens per chunk
    • 10–20% overlap
  • Chunk by semantic boundaries when possible:
    • headings
    • paragraphs
    • sections
    • code blocks kept intact

2) Make sure you’re embedding the right text

For documents, embed:

  • the chunk text itself
  • optionally prepend metadata like:
    • title
    • section heading
    • source type

Example:

  • Title: Refund Policy
  • Section: Eligibility
  • Text: Customers may request...

This often improves retrieval a lot.

3) Improve the query

User queries are often short and vague. If possible:

  • rewrite queries into more explicit search intents
  • expand acronyms
  • add context from conversation history

Example:

  • Bad: “limits”
  • Better: “What are the API rate limits for GPT-4o requests?”

4) Check your similarity metric

Most embedding retrieval works best with:

  • cosine similarity
  • or normalized dot product

If your index uses raw Euclidean distance or unnormalized vectors incorrectly, results can degrade.

5) Use a better retrieval pipeline

Pure vector search is often not enough.

A strong setup is:

  1. Vector search for top 20–50 chunks
  2. Keyword/BM25 hybrid search
  3. Rerank the candidates with a cross-encoder or LLM reranker

This is especially helpful when exact terms matter.

6) Tune top-k and thresholds

If you only retrieve top 3, you may miss the right chunk.

Try:

  • top-k = 10, 20, or 50
  • then rerank
  • if you use a similarity cutoff, make sure it’s not too strict

7) Inspect failure modes

Look at examples where retrieval fails and classify them:

  • wrong chunking?
  • query too vague?
  • term mismatch?
  • document contains answer but in a different section?
  • multiple near-duplicate chunks?

This often makes the fix obvious.

8) Consider using hybrid search

OpenAI embeddings are semantic, but they don’t always handle:

  • exact names
  • IDs
  • error codes
  • product SKUs
  • legal citations

Hybrid search helps a lot here.

9) Re-embed with a strong model

If you’re using an older or weaker embedding model, upgrade. For OpenAI, use a current embedding model rather than legacy ones.

10) Measure retrieval quality directly

Don’t judge by generation quality alone. Measure:

  • Recall@k
  • MRR
  • precision@k

If the relevant chunk isn’t in the top-k, the issue is retrieval, not the LLM.


A good baseline architecture

If you want a reliable setup:

  • Chunk at 300–400 tokens with overlap
  • Embed chunk text + title/heading
  • Normalize vectors
  • Retrieve top 20 using cosine similarity
  • Add BM25 hybrid retrieval
  • Rerank top 20 with a reranker
  • Pass top 3–5 chunks to the LLM

If you want, I can help debug this concretely

If you share any of the following, I can tell you what’s likely going wrong:

  • your chunk size / overlap
  • a few example documents
  • a sample query and the retrieved results
  • which embedding model and vector DB you’re using
  • whether you use cosine, dot product, or Euclidean distance

If you want, I can also give you a step-by-step retrieval debugging checklist or help you redesign your RAG pipeline.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.