Prompt
How do I build a RAG app with a vector database?
Latest observation
A RAG app (Retrieval-Augmented Generation) with a vector database usually has 4 parts:
- Ingest documents
- Chunk and embed them
- Store embeddings in a vector DB
- Retrieve relevant chunks at query time and send them to an LLM
Here’s the practical blueprint.
1) Basic architecture
Offline / indexing flow
- Load documents: PDFs, docs, web pages, markdown, database rows, etc.
- Clean and split into chunks.
- Create embeddings for each chunk.
- Store:
- vector embedding
- chunk text
- metadata like source, title, page, timestamp, permissions
Online / query flow
- User asks a question.
- Embed the question with the same embedding model.
- Search the vector DB for top-k similar chunks.
- Optionally rerank results.
- Put retrieved context into the LLM prompt.
- LLM generates answer grounded in retrieved text.
2) Choose your stack
Common options:
Vector databases
- Pinecone: managed, easy scaling
- Weaviate: open source + managed, feature-rich
- Qdrant: fast, great filtering
- Milvus: strong for large-scale workloads
- Chroma: simple local/dev-friendly
- pgvector: if you want to keep everything in Postgres
Embeddings
- OpenAI text embeddings
- SentenceTransformers / Hugging Face models
- Cohere embeddings
- Voyage embeddings
LLMs
- OpenAI GPT models
- Anthropic Claude
- Mistral, Llama, etc.
Frameworks
- LlamaIndex
- LangChain
- Haystack
- Or build it yourself for more control
3) Chunking strategy
This matters a lot.
Good defaults:
- Chunk size: 300–800 tokens
- Overlap: 50–150 tokens
Tips:
- Split by headings/sections when possible.
- Keep tables/code blocks intact if they matter.
- Store metadata for each chunk:
- document id
- section title
- page number
- URL/source
- access control info
Bad chunking leads to poor retrieval even if embeddings are good.
4) Indexing pipeline example
Pseudo-steps:
for each document:
text = extract_text(document)
chunks = split_into_chunks(text)
for each chunk:
embedding = embed(chunk)
vector_db.upsert(
id=chunk_id,
vector=embedding,
payload={
"text": chunk,
"source": document स्रोत,
"page": page_number,
"metadata": ...
}
)
5) Query pipeline example
user_question -> embed(question)
-> vector_db.search(top_k=5)
-> rerank results (optional)
-> build prompt with context
-> LLM answers
Example prompt structure:
You are a helpful assistant. Answer only using the provided context.
If the context does not contain the answer, say you don't know.
Context:
[1] ...
[2] ...
[3] ...
Question: ...
Answer:
6) Minimal Python example with a vector DB
Below is a simplified example using a hypothetical vector DB client.
from openai import OpenAI
client = OpenAI()
def embed_text(text):
resp = client.embeddings.create(
model="text-embedding-3-small",
input=text
)
return resp.data[0].embedding
def index_chunk(vector_db, chunk_id, text, metadata):
vector = embed_text(text)
vector_db.upsert(
id=chunk_id,
vector=vector,
payload={"text": text, **metadata}
)
def answer_question(vector_db, question):
q_vec = embed_text(question)
matches = vector_db.search(vector=q_vec, top_k=5)
context = "\n\n".join(
f"[{i+1}] {m.payload['text']}" for i, m in enumerate(matches)
)
prompt = f"""Answer using only the context below.
Context:
{context}
Question: {question}
"""
resp = client.responses.create(
model="gpt-4.1-mini",
input=prompt
)
return resp.output_text
7) Improve retrieval quality
Once the basic app works, improve it with:
A. Better chunking
- Use semantic chunking
- Chunk by headings
- Keep related sections together
B. Metadata filtering
Filter by:
- tenant/user
- document type
- time range
- language
- access permissions
Example: only search docs for the current user.
C. Hybrid search
Combine:
- dense vector search
- keyword/BM25 search
This helps with:
- exact names
- acronyms
- IDs
- rare terms
D. Reranking
Use a cross-encoder or LLM-based reranker to reorder top results for relevance.
E. Context compression
If retrieved chunks are long, summarize or extract only relevant sentences before sending to the LLM.
8) Production concerns
Latency
- Precompute embeddings
- Cache frequent queries
- Use top-k small enough
- Consider reranking only when necessary
Security
- Enforce per-user or per-tenant filters
- Don’t retrieve unauthorized documents
- Log carefully to avoid leaking sensitive text
Freshness
- Re-index changed documents
- Use incremental updates
- Track document versions
Evaluation
Measure:
- retrieval accuracy
- answer faithfulness
- hallucination rate
- latency
- cost
You can test with:
- manually labeled question-answer pairs
- synthetic queries
- retrieval metrics like recall@k
9) Recommended MVP stack
If you want the simplest path:
- Python
- LlamaIndex or LangChain
- Chroma or Qdrant
- OpenAI embeddings
- GPT model for generation
This gets you a working prototype quickly.
10) A good MVP plan
- Pick 1 document source
- Chunk and embed documents
- Store in vector DB
- Build a simple question-answer endpoint
- Add metadata filters
- Add reranking
- Add auth/permissions
- Evaluate with test questions
- Deploy
If you want, I can also give you:
- a full working code example
- a LangChain version
- a LlamaIndex version
- or a FastAPI + Qdrant + OpenAI reference implementation.