Prompt
How do I set up vector search for a RAG app?
Latest observation
To set up vector search for a RAG app, you need four main pieces:
- Document ingestion
- Chunking + embedding
- Vector storage + retrieval
- LLM generation with retrieved context
Here’s a practical setup.
1) Ingest your data
Collect the documents you want your app to answer questions from, such as:
- PDFs
- web pages
- Markdown / text files
- database records
- support tickets
- internal docs
You’ll want to normalize them into plain text with metadata like:
- source
- title
- URL
- date
- section/page
- access permissions
Metadata matters because it helps with filtering, citations, and debugging.
2) Chunk the documents
Don’t embed huge documents as one block. Split them into smaller chunks.
Typical chunking:
- Chunk size: 300–1,000 tokens
- Overlap: 50–150 tokens
Good chunking practices:
- Keep chunks semantically coherent
- Prefer splitting by headings/paragraphs over arbitrary token cuts
- Preserve metadata for each chunk
- Avoid chunks that are too tiny or too large
Example:
Document -> sections -> paragraphs -> chunks
3) Create embeddings
Convert each chunk into a vector using an embedding model.
Popular embedding options:
- OpenAI embeddings
- Cohere embeddings
- Sentence Transformers
- Voyage / Jina / Mistral embeddings
For each chunk, store:
- chunk text
- embedding vector
- metadata
Example flow:
embedding = embed_model.encode(chunk_text)
4) Store vectors in a vector database
You need a place to index and search embeddings.
Popular vector stores:
- FAISS — local, simple, fast
- Pinecone — managed, scalable
- Weaviate — vector DB + metadata filtering
- Milvus — open-source, scalable
- Qdrant — strong metadata filtering, easy to use
- Chroma — good for prototypes
When storing, use:
- vector
- chunk text
- metadata
- unique ID
5) Retrieve relevant chunks at query time
When a user asks a question:
- Embed the query
- Search the vector DB for nearest chunks
- Return top-k results
Usually:
- top-k = 3 to 10
- optionally use metadata filters
- optionally rerank results with a cross-encoder or LLM reranker
Example:
query_vec = embed_model.encode(user_query)
results = vector_db.search(query_vec, top_k=5)
6) Build the RAG prompt
Pass the retrieved chunks into the LLM as context, along with the user question.
A simple prompt structure:
You are a helpful assistant. Answer only using the context below.
Context:
[chunk 1]
[chunk 2]
[chunk 3]
Question:
[user question]
Best practices:
- Include source citations
- Tell the model to say “I don’t know” if context is insufficient
- Keep the context window within model limits
- Use only the most relevant chunks
7) Add reranking and filtering
Basic vector search is often good, but for better quality:
Reranking
After retrieving top 10–20 chunks, rerank them with:
- a cross-encoder
- an LLM-based reranker
- relevance scoring
Filtering
Use metadata filters:
- document type
- date range
- user permissions
- language
- tenant/customer ID
This improves precision and security.
8) Evaluate and tune
Measure retrieval quality and answer quality.
Track:
- recall@k
- precision@k
- answer correctness
- citation accuracy
- hallucination rate
Tune:
- chunk size
- overlap
- embedding model
- top-k
- reranking
- metadata filters
Simple architecture
Documents
-> text extraction
-> chunking
-> embedding
-> vector DB
User query
-> query embedding
-> vector search
-> rerank
-> prompt assembly
-> LLM answer
Minimal example with FAISS-like flow
from sentence_transformers import SentenceTransformer
import faiss
import numpy as np
model = SentenceTransformer("all-MiniLM-L6-v2")
chunks = [
"RAG combines retrieval and generation.",
"Vector search finds semantically similar text.",
"Chunking improves retrieval quality."
]
embeddings = model.encode(chunks, normalize_embeddings=True)
embeddings = np.array(embeddings).astype("float32")
index = faiss.IndexFlatIP(embeddings.shape[1])
index.add(embeddings)
query = "How does vector search help RAG?"
q_emb = model.encode([query], normalize_embeddings=True).astype("float32")
scores, ids = index.search(q_emb, k=2)
for i in ids[0]:
print(chunks[i])
Recommended stack for a production RAG app
A common setup:
- Parsing: Unstructured, Apache Tika, custom loaders
- Chunking: custom or LangChain/LlamaIndex
- Embeddings: OpenAI / Cohere / Voyage / Sentence Transformers
- Vector DB: Qdrant, Pinecone, Weaviate, or Milvus
- Reranking: Cohere rerank, bge-reranker, or custom
- LLM: GPT-4.1 / Claude / Gemini / local model
- Orchestration: LangChain, LlamaIndex, or your own code
Key best practices
- Use high-quality chunking
- Store good metadata
- Tune top-k and chunk size
- Add reranking for better relevance
- Filter by permissions if needed
- Log retrieval results for debugging
- Keep embeddings and vector DB in sync with source docs
If you want, I can also give you:
- a production-ready architecture diagram,
- a Python implementation with Qdrant or Pinecone, or
- a step-by-step setup using LangChain/LlamaIndex.