Prompt

I’m unhappy with Cohere embeddings on my support corpus

Artificial Intelligence / AI Search1 observationLast seen Oct 10, 2026

Latest observation

Oct 10, 2026OpenAI APIWeb search: off

Sorry to hear that. “Unhappy” with embeddings can mean a few different failure modes, and the fix depends on which one you’re seeing.

Common issues with support corpora embeddings:

  1. Retrieval misses obvious docs

    • Queries use different wording than tickets/docs.
    • The embedding model isn’t handling domain-specific language well.
  2. Too many near-duplicate results

    • Support articles may be repetitive, causing results to collapse around generic text.
  3. Bad chunking

    • Chunks are too large, too small, or split across helpful context boundaries.
  4. Poor semantic separation

    • Similar but distinct issues (billing vs login vs device activation) cluster too closely.
  5. Short query problem

    • User queries are terse (“can’t sign in”), which can be hard to embed well without query expansion or reranking.
  6. Language / jargon mismatch

    • Product names, error codes, and internal terminology can confuse general-purpose embeddings.

What usually helps:

  • Evaluate with a small labeled set

    • Pick 50–200 real queries and mark the correct docs.
    • Measure Recall@k / MRR before changing anything.
  • Use better chunking

    • Try chunk sizes around 200–500 tokens with overlap.
    • Keep titles, headings, and error codes attached to chunks.
  • Add reranking

    • Use embeddings for candidate retrieval, then a reranker for final ordering.
    • This often gives a noticeable lift on support search.
  • Normalize and enrich text

    • Include ticket subject + body + product + tags + FAQ title.
    • Expand acronyms and preserve error codes.
  • Try a different embedding model

    • Especially one tuned for retrieval and short queries.
    • Cohere may be fine, but your corpus/query mix may fit another model better.
  • Hybrid search

    • Combine lexical search (BM25) with embeddings.
    • Support content often benefits a lot from exact-match signals like error codes and product names.

If you want, I can help you diagnose this more concretely. Send me any of:

  • 3–5 example queries
  • the top results you’re getting
  • the result you expected
  • how you chunked the docs
  • which Cohere embedding model you’re using

Then I can suggest specific changes.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.