Prompt

How do I build a document Q&A feature with an AI API?

Artificial Intelligence / AI Platforms1 observationLast seen Oct 8, 2026

Latest observation

Oct 8, 2026OpenAI APIWeb search: off

A solid document Q&A feature usually looks like this:

  1. Ingest documents
  2. Split them into chunks
  3. Create embeddings for each chunk
  4. Store embeddings in a vector database
  5. On each question, retrieve the most relevant chunks
  6. Send those chunks + the question to the AI model
  7. Return an answer, ideally with citations

This is commonly called RAG: Retrieval-Augmented Generation.


1) Basic architecture

Offline / background pipeline

  • Upload PDF, DOCX, HTML, text, etc.
  • Extract clean text
  • Split into chunks
  • Generate embeddings for each chunk using an embeddings model
  • Store:
    • chunk text
    • embedding vector
    • metadata like document ID, page number, section title, permissions

Query-time pipeline

  • User asks a question
  • Embed the question
  • Search vector DB for top-k similar chunks
  • Optionally rerank results
  • Provide retrieved chunks to a chat/completions model
  • Ask the model to answer only using those chunks
  • Return answer + sources

2) Choose your stack

Common options

  • LLM / API: OpenAI API or another AI provider
  • Embeddings: text-embedding model
  • Vector DB:
    • Pinecone
    • Weaviate
    • Qdrant
    • Milvus
    • pgvector in Postgres
  • Backend: Python/FastAPI, Node/Express, etc.
  • Document parsing:
    • PDFs: pypdf, pdfplumber, unstructured
    • DOCX: python-docx
    • OCR if scanned docs: Tesseract, cloud OCR

If you want something simple and cheap to start, Postgres + pgvector is often a great choice.


3) Chunking strategy

Good chunking matters a lot.

Practical defaults

  • Chunk size: 300–800 tokens
  • Overlap: 50–150 tokens
  • Keep chunks semantically coherent when possible:
    • split by headings/paragraphs
    • avoid breaking tables in the middle
    • preserve page numbers and section titles

Why chunking matters

  • Too large: retrieval gets noisy and expensive
  • Too small: context lacks enough information

4) Store useful metadata

For each chunk store:

  • document_id
  • chunk_id
  • text
  • embedding
  • page_number
  • section_title
  • source_file_name
  • user_id / tenant_id for access control

This makes citation, filtering, and permissions much easier.


5) Retrieval flow

At query time:

  1. Embed the question
  2. Vector search top 5–20 chunks
  3. Filter by access permissions or document scope
  4. Optionally rerank with a cross-encoder or LLM
  5. Build prompt with retrieved text
  6. Generate answer

Tip

Use retrieval filters like:

  • same customer/account
  • only documents user has access to
  • date ranges
  • document type

6) Prompting the model

A strong prompt should:

  • instruct the model to use only provided context
  • say “I don’t know” if context is insufficient
  • request citations
  • keep answer concise or detailed based on your UX

Example prompt structure:

You are a helpful assistant answering questions about documents.
Use only the provided context to answer.
If the answer is not in the context, say you don't know.

Context:
[chunk 1 text]
[chunk 2 text]
[chunk 3 text]

Question:
What is the refund policy?

Answer with citations to the chunks you used.

7) Example implementation outline

Ingestion

def ingest_document(doc_id, text):
    chunks = split_into_chunks(text)
    for i, chunk in enumerate(chunks):
        embedding = embed(chunk)
        save_chunk(
            document_id=doc_id,
            chunk_id=i,
            text=chunk,
            embedding=embedding
        )

Query

def answer_question(question, user_id):
    q_embedding = embed(question)
    chunks = search_similar_chunks(q_embedding, top_k=5, user_id=user_id)

    context = "\n\n".join(
        f"[Source {i+1}] {chunk['text']}"
        for i, chunk in enumerate(chunks)
    )

    prompt = f"""
You are a helpful assistant answering questions about documents.
Use only the context below. If the answer is not present, say you don't know.

Context:
{context}

Question: {question}
"""

    return generate_answer(prompt)

8) How to improve quality

Add reranking

Vector search is good for recall, but not always precision.
A reranker can reorder the retrieved chunks by relevance.

Use hybrid search

Combine:

  • vector similarity
  • keyword/BM25 search

This helps with:

  • names
  • IDs
  • exact terminology
  • legal/technical terms

Add citations

Return:

  • document name
  • page number
  • chunk text snippet

Users trust answers much more with sources.

Summarize long answers

For very large context, do:

  • retrieve
  • synthesize short evidence summaries
  • answer from summaries

Handle follow-up questions

Maintain chat history, but don’t blindly stuff all history into the prompt.
Instead:

  • keep short conversation memory
  • retrieve from docs again for each new question

9) Security and access control

Important for enterprise or multi-user apps:

  • enforce permissions before retrieval
  • do not retrieve unauthorized chunks
  • separate embeddings/indexes by tenant if needed
  • log access for compliance

10) Common pitfalls

  • Bad parsing from PDFs, especially scanned documents
  • No metadata, making citations hard
  • Chunks too large/small
  • No access control
  • Using only vector search when keyword search is also needed
  • Letting model hallucinate by not constraining it with retrieved context
  • Not evaluating with real questions

11) Evaluation

Test your system with:

  • known question-answer pairs
  • answer exactness
  • citation correctness
  • retrieval recall
  • latency
  • hallucination rate

Track metrics like:

  • top-k recall
  • answer accuracy
  • source attribution accuracy
  • “I don’t know” correctness

12) A minimal production-ready flow

If you want the simplest reliable version:

  • Parse docs
  • Chunk by paragraph/heading
  • Embed with an embeddings API
  • Store in pgvector
  • Retrieve top 8 chunks
  • Rerank top 8 to top 3
  • Prompt model with only those 3 chunks
  • Return answer + citations
  • Log everything for evaluation

13) If you want to build it quickly with an AI API

Typical API usage pattern:

  • Embeddings endpoint for document chunks and questions
  • Chat/completions endpoint for answering
  • Optional file/search tools if your provider offers built-in document search

If your AI API supports built-in retrieval or file search, that can reduce infrastructure. But for more control, many teams still implement their own RAG stack.


If you want, I can give you one of these next:

  1. a Python example using OpenAI API + pgvector
  2. a Node.js example
  3. a system design diagram
  4. a fully working minimal RAG app tutorial

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.