Prompt

How do I build a RAG app with a vector database?

Technology / Databases1 observationLast seen Oct 7, 2026

Latest observation

Oct 7, 2026OpenAI APIWeb search: off

A RAG app (Retrieval-Augmented Generation) with a vector database usually has 4 parts:

  1. Ingest documents
  2. Chunk and embed them
  3. Store embeddings in a vector DB
  4. Retrieve relevant chunks at query time and send them to an LLM

Here’s the practical blueprint.


1) Basic architecture

Offline / indexing flow

  • Load documents: PDFs, docs, web pages, markdown, database rows, etc.
  • Clean and split into chunks.
  • Create embeddings for each chunk.
  • Store:
    • vector embedding
    • chunk text
    • metadata like source, title, page, timestamp, permissions

Online / query flow

  • User asks a question.
  • Embed the question with the same embedding model.
  • Search the vector DB for top-k similar chunks.
  • Optionally rerank results.
  • Put retrieved context into the LLM prompt.
  • LLM generates answer grounded in retrieved text.

2) Choose your stack

Common options:

Vector databases

  • Pinecone: managed, easy scaling
  • Weaviate: open source + managed, feature-rich
  • Qdrant: fast, great filtering
  • Milvus: strong for large-scale workloads
  • Chroma: simple local/dev-friendly
  • pgvector: if you want to keep everything in Postgres

Embeddings

  • OpenAI text embeddings
  • SentenceTransformers / Hugging Face models
  • Cohere embeddings
  • Voyage embeddings

LLMs

  • OpenAI GPT models
  • Anthropic Claude
  • Mistral, Llama, etc.

Frameworks

  • LlamaIndex
  • LangChain
  • Haystack
  • Or build it yourself for more control

3) Chunking strategy

This matters a lot.

Good defaults:

  • Chunk size: 300–800 tokens
  • Overlap: 50–150 tokens

Tips:

  • Split by headings/sections when possible.
  • Keep tables/code blocks intact if they matter.
  • Store metadata for each chunk:
    • document id
    • section title
    • page number
    • URL/source
    • access control info

Bad chunking leads to poor retrieval even if embeddings are good.


4) Indexing pipeline example

Pseudo-steps:

for each document:
    text = extract_text(document)
    chunks = split_into_chunks(text)
    for each chunk:
        embedding = embed(chunk)
        vector_db.upsert(
            id=chunk_id,
            vector=embedding,
            payload={
                "text": chunk,
                "source": document स्रोत,
                "page": page_number,
                "metadata": ...
            }
        )

5) Query pipeline example

user_question -> embed(question)
              -> vector_db.search(top_k=5)
              -> rerank results (optional)
              -> build prompt with context
              -> LLM answers

Example prompt structure:

You are a helpful assistant. Answer only using the provided context.
If the context does not contain the answer, say you don't know.

Context:
[1] ...
[2] ...
[3] ...

Question: ...
Answer:

6) Minimal Python example with a vector DB

Below is a simplified example using a hypothetical vector DB client.

from openai import OpenAI

client = OpenAI()

def embed_text(text):
    resp = client.embeddings.create(
        model="text-embedding-3-small",
        input=text
    )
    return resp.data[0].embedding

def index_chunk(vector_db, chunk_id, text, metadata):
    vector = embed_text(text)
    vector_db.upsert(
        id=chunk_id,
        vector=vector,
        payload={"text": text, **metadata}
    )

def answer_question(vector_db, question):
    q_vec = embed_text(question)
    matches = vector_db.search(vector=q_vec, top_k=5)

    context = "\n\n".join(
        f"[{i+1}] {m.payload['text']}" for i, m in enumerate(matches)
    )

    prompt = f"""Answer using only the context below.

Context:
{context}

Question: {question}
"""

    resp = client.responses.create(
        model="gpt-4.1-mini",
        input=prompt
    )
    return resp.output_text

7) Improve retrieval quality

Once the basic app works, improve it with:

A. Better chunking

  • Use semantic chunking
  • Chunk by headings
  • Keep related sections together

B. Metadata filtering

Filter by:

  • tenant/user
  • document type
  • time range
  • language
  • access permissions

Example: only search docs for the current user.

C. Hybrid search

Combine:

  • dense vector search
  • keyword/BM25 search

This helps with:

  • exact names
  • acronyms
  • IDs
  • rare terms

D. Reranking

Use a cross-encoder or LLM-based reranker to reorder top results for relevance.

E. Context compression

If retrieved chunks are long, summarize or extract only relevant sentences before sending to the LLM.


8) Production concerns

Latency

  • Precompute embeddings
  • Cache frequent queries
  • Use top-k small enough
  • Consider reranking only when necessary

Security

  • Enforce per-user or per-tenant filters
  • Don’t retrieve unauthorized documents
  • Log carefully to avoid leaking sensitive text

Freshness

  • Re-index changed documents
  • Use incremental updates
  • Track document versions

Evaluation

Measure:

  • retrieval accuracy
  • answer faithfulness
  • hallucination rate
  • latency
  • cost

You can test with:

  • manually labeled question-answer pairs
  • synthetic queries
  • retrieval metrics like recall@k

9) Recommended MVP stack

If you want the simplest path:

  • Python
  • LlamaIndex or LangChain
  • Chroma or Qdrant
  • OpenAI embeddings
  • GPT model for generation

This gets you a working prototype quickly.


10) A good MVP plan

  1. Pick 1 document source
  2. Chunk and embed documents
  3. Store in vector DB
  4. Build a simple question-answer endpoint
  5. Add metadata filters
  6. Add reranking
  7. Add auth/permissions
  8. Evaluate with test questions
  9. Deploy

If you want, I can also give you:

  • a full working code example
  • a LangChain version
  • a LlamaIndex version
  • or a FastAPI + Qdrant + OpenAI reference implementation.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.