Prompt
How do I use embeddings for RAG retrieval over internal docs?
Latest observation
A simple RAG setup for internal docs looks like this:
1) Ingest and chunk the docs
Break documents into smaller passages so retrieval can find the right section.
- Typical chunk size: 300–800 tokens
- Overlap: 50–150 tokens
- Keep metadata:
doc_idtitlesectiondateaccess_control / department
Why: embeddings work best on semantically coherent chunks, not whole long docs.
2) Create embeddings for each chunk
Use an embedding model to convert each chunk into a vector.
Store:
- the chunk text
- its vector
- metadata
Example conceptually:
embedding = embed_model.encode(chunk_text)
vector_db.add(
id=chunk_id,
vector=embedding,
text=chunk_text,
metadata=...
)
3) Put vectors in a vector database or search index
Common options:
- Pinecone
- Weaviate
- Qdrant
- Milvus
- FAISS (local)
- Elasticsearch/OpenSearch with vector search
You’ll query this store later with an embedded user question.
4) Embed the user query at retrieval time
When a user asks a question:
- embed the query
- retrieve nearest chunks by cosine similarity / dot product
- optionally filter by metadata like department, document type, or permissions
q_vec = embed_model.encode(user_query)
results = vector_db.search(q_vec, top_k=5, filter={"team": "finance"})
5) Send retrieved chunks to the LLM
Take the top results and place them in the prompt as context.
Example structure:
You are answering using the following internal documents:
[chunk 1]
[chunk 2]
[chunk 3]
Question: ...
Answer using only the provided context.
The LLM then generates an answer grounded in those retrieved passages.
Practical best practices
Chunking
- Split by headings/paragraphs if possible
- Avoid chopping in the middle of lists or tables
- Keep chunks self-contained
- If docs are highly structured, store section titles in the chunk text
Retrieval quality
- Use top_k 5–20
- Consider hybrid search:
- embeddings for semantic match
- keyword/BM25 for exact term match
- Consider reranking with a cross-encoder or LLM to improve final chunk selection
Metadata filtering
Very important for internal docs:
- department
- access level
- doc freshness
- document type
- language
This prevents retrieving irrelevant or unauthorized content.
Freshness
If docs change often:
- re-embed changed chunks only
- keep versioning
- periodically reindex
Evaluation
Test with real questions and measure:
- retrieval recall
- answer faithfulness
- citation correctness
- latency
Minimal architecture
Docs -> chunking -> embeddings -> vector DB
User question -> query embedding -> retrieve top chunks -> LLM answer
Common pitfalls
- Chunks too large: retrieval becomes vague
- Chunks too small: missing context
- No metadata filtering: wrong or unauthorized docs
- Using only vector search: misses exact names, IDs, acronyms
- No reranking: top results may be semantically close but not best
Example workflow
If your internal docs include policies, runbooks, and FAQs:
- Split each into chunks by heading
- Embed each chunk
- Store in vector DB with metadata
- At query time:
- filter by access permissions
- retrieve top 10
- rerank to top 3
- give those 3 to the LLM
- Return answer with citations to source docs
If you want, I can also show:
- a Python code example using FAISS or Qdrant
- a production architecture
- a best-practice chunking strategy for PDFs/Notion/Confluence docs