Prompt

I'm building a chatbot over company docs and need an embeddings strategy

Artificial Intelligence / AI Search1 observationLast seen Oct 10, 2026

Latest observation

Oct 10, 2026OpenAI APIWeb search: off

Here’s a solid embeddings strategy for a chatbot over company docs, optimized for reliability, retrieval quality, and maintainability.

1) Start with the right retrieval architecture

For company docs, the best default is usually:

  • Chunk documents
  • Embed chunks
  • Store in a vector database
  • Retrieve top-k chunks per query
  • Optionally re-rank
  • Generate answer grounded in retrieved chunks

This is the standard RAG setup and works well for policies, manuals, knowledge bases, PDFs, wiki pages, tickets, etc.


2) Choose an embedding model based on your constraints

Pick an embedding model with strong semantic retrieval performance, good multilingual support if needed, and reasonable cost.

Good options

  • OpenAI text-embedding-3-large: best quality, more expensive
  • OpenAI text-embedding-3-small: cheaper, strong baseline
  • Cohere Embed v3
  • Voyage embeddings
  • Sentence-transformers / BGE / E5 if you want self-hosted

Rule of thumb

  • If accuracy matters most: use a top-tier hosted model
  • If cost/latency matters most: use a smaller hosted model
  • If data residency or control matters: self-hosted embeddings

3) Chunking strategy matters as much as the model

Bad chunking can ruin retrieval even with great embeddings.

Recommended chunking

  • Chunk by semantic structure when possible:
    • headings
    • sections
    • paragraphs
    • bullet lists
  • Use token-based chunks with overlap as fallback

Typical settings

  • Chunk size: 300–800 tokens
  • Overlap: 50–150 tokens

Best practices

  • Keep chunks focused on one topic
  • Avoid cutting tables or lists in the middle
  • Preserve metadata:
    • document title
    • section heading
    • source URL/path
    • last updated date
    • access permissions

4) Use metadata aggressively

Metadata improves filtering, ranking, and trust.

Store for each chunk:

  • doc_id
  • chunk_id
  • title
  • heading path
  • source
  • department
  • created_at / updated_at
  • version
  • ACL / permissions
  • document type

Then use metadata filters like:

  • only show HR docs to HR users
  • only current policy versions
  • only docs from a specific region/business unit

5) Use hybrid retrieval, not just vector search

For enterprise docs, hybrid search is usually better than vector-only.

Combine:

  • Dense embeddings for semantic matching
  • Sparse retrieval like BM25 for exact terms, acronyms, part numbers, policy names

This helps when users ask:

  • “What is the PTO policy?”
  • “How do I reset Okta?”
  • “What is SOX control 3.2?”
  • “Where is the Q4 revenue deck?”

Exact terms and jargon often matter a lot.


6) Add a reranker if quality matters

A reranker can significantly improve final retrieval quality.

Pipeline:

  1. Retrieve top 20–50 chunks using hybrid search
  2. Re-rank them with a cross-encoder/reranker
  3. Send top 3–8 chunks to the LLM

This reduces irrelevant context and improves answer accuracy.


7) Consider multiple embeddings per document type

Different document types benefit from different handling:

  • Policy docs: chunk by section, preserve hierarchy
  • FAQs: embed question-answer pairs
  • Tables: convert rows to text or row-level chunks
  • Slides: often need slide-level chunks plus speaker notes if available
  • Tickets / chat logs: thread-level chunks or summarized chunks

8) Plan for updates and versioning

Company docs change constantly.

You want:

  • incremental re-embedding for changed docs only
  • versioned chunks
  • soft deletion of obsolete docs
  • re-indexing pipeline triggered by doc updates

Also consider a TTL or freshness boost:

  • prefer newer docs when content is similar
  • down-rank stale documents

9) Evaluate retrieval, not just answer quality

Measure embeddings strategy using retrieval metrics.

Useful metrics

  • Recall@k
  • MRR
  • nDCG
  • Answer faithfulness
  • Groundedness
  • Citation accuracy

Build a test set of real user questions with expected source docs.

What to test

  • exact terminology queries
  • paraphrases
  • ambiguous queries
  • multi-hop questions
  • acronym-heavy questions
  • stale vs current policy distinctions

10) Practical recommended default stack

If you want a strong, simple starting point:

  • Chunking: 500 tokens with 100 token overlap, structure-aware
  • Embeddings: a strong hosted embedding model
  • Index: vector DB with metadata filtering
  • Retrieval: hybrid dense + BM25
  • Reranking: yes
  • Generation: answer only from retrieved context, with citations
  • Freshness: prefer latest approved docs
  • Access control: enforce at retrieval time

11) Common pitfalls to avoid

  • Using huge chunks that mix multiple topics
  • Ignoring metadata and document versioning
  • Vector-only search in enterprise settings
  • Not handling acronyms and exact terms
  • Re-embedding everything on every change
  • Not testing retrieval on real user questions
  • Letting the model answer without citations or grounding

12) If you want the fastest path to production

Do this:

  1. Ingest docs
  2. Parse structure
  3. Chunk by section/paragraph
  4. Embed chunks
  5. Store in vector DB with metadata
  6. Add BM25 hybrid search
  7. Add reranker
  8. Build evaluation set
  9. Iterate on chunking and retrieval

If you want, I can also give you:

  1. a specific recommended stack for your budget,
  2. a chunking recipe for PDFs/wiki/HTML, or
  3. a reference architecture diagram for the whole chatbot pipeline.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.