Prompt
How do I build a document Q&A feature with an AI API?
Latest observation
A solid document Q&A feature usually looks like this:
- Ingest documents
- Split them into chunks
- Create embeddings for each chunk
- Store embeddings in a vector database
- On each question, retrieve the most relevant chunks
- Send those chunks + the question to the AI model
- Return an answer, ideally with citations
This is commonly called RAG: Retrieval-Augmented Generation.
1) Basic architecture
Offline / background pipeline
- Upload PDF, DOCX, HTML, text, etc.
- Extract clean text
- Split into chunks
- Generate embeddings for each chunk using an embeddings model
- Store:
- chunk text
- embedding vector
- metadata like document ID, page number, section title, permissions
Query-time pipeline
- User asks a question
- Embed the question
- Search vector DB for top-k similar chunks
- Optionally rerank results
- Provide retrieved chunks to a chat/completions model
- Ask the model to answer only using those chunks
- Return answer + sources
2) Choose your stack
Common options
- LLM / API: OpenAI API or another AI provider
- Embeddings: text-embedding model
- Vector DB:
- Pinecone
- Weaviate
- Qdrant
- Milvus
- pgvector in Postgres
- Backend: Python/FastAPI, Node/Express, etc.
- Document parsing:
- PDFs: pypdf, pdfplumber, unstructured
- DOCX: python-docx
- OCR if scanned docs: Tesseract, cloud OCR
If you want something simple and cheap to start, Postgres + pgvector is often a great choice.
3) Chunking strategy
Good chunking matters a lot.
Practical defaults
- Chunk size: 300–800 tokens
- Overlap: 50–150 tokens
- Keep chunks semantically coherent when possible:
- split by headings/paragraphs
- avoid breaking tables in the middle
- preserve page numbers and section titles
Why chunking matters
- Too large: retrieval gets noisy and expensive
- Too small: context lacks enough information
4) Store useful metadata
For each chunk store:
document_idchunk_idtextembeddingpage_numbersection_titlesource_file_nameuser_id/tenant_idfor access control
This makes citation, filtering, and permissions much easier.
5) Retrieval flow
At query time:
- Embed the question
- Vector search top 5–20 chunks
- Filter by access permissions or document scope
- Optionally rerank with a cross-encoder or LLM
- Build prompt with retrieved text
- Generate answer
Tip
Use retrieval filters like:
- same customer/account
- only documents user has access to
- date ranges
- document type
6) Prompting the model
A strong prompt should:
- instruct the model to use only provided context
- say “I don’t know” if context is insufficient
- request citations
- keep answer concise or detailed based on your UX
Example prompt structure:
You are a helpful assistant answering questions about documents.
Use only the provided context to answer.
If the answer is not in the context, say you don't know.
Context:
[chunk 1 text]
[chunk 2 text]
[chunk 3 text]
Question:
What is the refund policy?
Answer with citations to the chunks you used.
7) Example implementation outline
Ingestion
def ingest_document(doc_id, text):
chunks = split_into_chunks(text)
for i, chunk in enumerate(chunks):
embedding = embed(chunk)
save_chunk(
document_id=doc_id,
chunk_id=i,
text=chunk,
embedding=embedding
)
Query
def answer_question(question, user_id):
q_embedding = embed(question)
chunks = search_similar_chunks(q_embedding, top_k=5, user_id=user_id)
context = "\n\n".join(
f"[Source {i+1}] {chunk['text']}"
for i, chunk in enumerate(chunks)
)
prompt = f"""
You are a helpful assistant answering questions about documents.
Use only the context below. If the answer is not present, say you don't know.
Context:
{context}
Question: {question}
"""
return generate_answer(prompt)
8) How to improve quality
Add reranking
Vector search is good for recall, but not always precision.
A reranker can reorder the retrieved chunks by relevance.
Use hybrid search
Combine:
- vector similarity
- keyword/BM25 search
This helps with:
- names
- IDs
- exact terminology
- legal/technical terms
Add citations
Return:
- document name
- page number
- chunk text snippet
Users trust answers much more with sources.
Summarize long answers
For very large context, do:
- retrieve
- synthesize short evidence summaries
- answer from summaries
Handle follow-up questions
Maintain chat history, but don’t blindly stuff all history into the prompt.
Instead:
- keep short conversation memory
- retrieve from docs again for each new question
9) Security and access control
Important for enterprise or multi-user apps:
- enforce permissions before retrieval
- do not retrieve unauthorized chunks
- separate embeddings/indexes by tenant if needed
- log access for compliance
10) Common pitfalls
- Bad parsing from PDFs, especially scanned documents
- No metadata, making citations hard
- Chunks too large/small
- No access control
- Using only vector search when keyword search is also needed
- Letting model hallucinate by not constraining it with retrieved context
- Not evaluating with real questions
11) Evaluation
Test your system with:
- known question-answer pairs
- answer exactness
- citation correctness
- retrieval recall
- latency
- hallucination rate
Track metrics like:
- top-k recall
- answer accuracy
- source attribution accuracy
- “I don’t know” correctness
12) A minimal production-ready flow
If you want the simplest reliable version:
- Parse docs
- Chunk by paragraph/heading
- Embed with an embeddings API
- Store in pgvector
- Retrieve top 8 chunks
- Rerank top 8 to top 3
- Prompt model with only those 3 chunks
- Return answer + citations
- Log everything for evaluation
13) If you want to build it quickly with an AI API
Typical API usage pattern:
- Embeddings endpoint for document chunks and questions
- Chat/completions endpoint for answering
- Optional file/search tools if your provider offers built-in document search
If your AI API supports built-in retrieval or file search, that can reduce infrastructure. But for more control, many teams still implement their own RAG stack.
If you want, I can give you one of these next:
- a Python example using OpenAI API + pgvector
- a Node.js example
- a system design diagram
- a fully working minimal RAG app tutorial