Prompt

How do I use embeddings for RAG retrieval over internal docs?

Artificial Intelligence / AI Search1 observationLast seen Oct 10, 2026

Latest observation

Oct 10, 2026OpenAI APIWeb search: off

A simple RAG setup for internal docs looks like this:

1) Ingest and chunk the docs

Break documents into smaller passages so retrieval can find the right section.

  • Typical chunk size: 300–800 tokens
  • Overlap: 50–150 tokens
  • Keep metadata:
    • doc_id
    • title
    • section
    • date
    • access_control / department

Why: embeddings work best on semantically coherent chunks, not whole long docs.


2) Create embeddings for each chunk

Use an embedding model to convert each chunk into a vector.

Store:

  • the chunk text
  • its vector
  • metadata

Example conceptually:

embedding = embed_model.encode(chunk_text)
vector_db.add(
    id=chunk_id,
    vector=embedding,
    text=chunk_text,
    metadata=...
)

3) Put vectors in a vector database or search index

Common options:

  • Pinecone
  • Weaviate
  • Qdrant
  • Milvus
  • FAISS (local)
  • Elasticsearch/OpenSearch with vector search

You’ll query this store later with an embedded user question.


4) Embed the user query at retrieval time

When a user asks a question:

  1. embed the query
  2. retrieve nearest chunks by cosine similarity / dot product
  3. optionally filter by metadata like department, document type, or permissions
q_vec = embed_model.encode(user_query)
results = vector_db.search(q_vec, top_k=5, filter={"team": "finance"})

5) Send retrieved chunks to the LLM

Take the top results and place them in the prompt as context.

Example structure:

You are answering using the following internal documents:
[chunk 1]
[chunk 2]
[chunk 3]

Question: ...
Answer using only the provided context.

The LLM then generates an answer grounded in those retrieved passages.


Practical best practices

Chunking

  • Split by headings/paragraphs if possible
  • Avoid chopping in the middle of lists or tables
  • Keep chunks self-contained
  • If docs are highly structured, store section titles in the chunk text

Retrieval quality

  • Use top_k 5–20
  • Consider hybrid search:
    • embeddings for semantic match
    • keyword/BM25 for exact term match
  • Consider reranking with a cross-encoder or LLM to improve final chunk selection

Metadata filtering

Very important for internal docs:

  • department
  • access level
  • doc freshness
  • document type
  • language

This prevents retrieving irrelevant or unauthorized content.

Freshness

If docs change often:

  • re-embed changed chunks only
  • keep versioning
  • periodically reindex

Evaluation

Test with real questions and measure:

  • retrieval recall
  • answer faithfulness
  • citation correctness
  • latency

Minimal architecture

Docs -> chunking -> embeddings -> vector DB
User question -> query embedding -> retrieve top chunks -> LLM answer

Common pitfalls

  • Chunks too large: retrieval becomes vague
  • Chunks too small: missing context
  • No metadata filtering: wrong or unauthorized docs
  • Using only vector search: misses exact names, IDs, acronyms
  • No reranking: top results may be semantically close but not best

Example workflow

If your internal docs include policies, runbooks, and FAQs:

  1. Split each into chunks by heading
  2. Embed each chunk
  3. Store in vector DB with metadata
  4. At query time:
    • filter by access permissions
    • retrieve top 10
    • rerank to top 3
    • give those 3 to the LLM
  5. Return answer with citations to source docs

If you want, I can also show:

  • a Python code example using FAISS or Qdrant
  • a production architecture
  • a best-practice chunking strategy for PDFs/Notion/Confluence docs

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.