Prompt
How do I build private Q&A over company documents?
Latest observation
To build private Q&A over company documents, you usually want a RAG system (Retrieval-Augmented Generation) rather than fine-tuning. That keeps the model private, grounded in your docs, and easier to update.
High-level architecture
-
Ingest documents
- Connect to sources: Google Drive, SharePoint, Confluence, Notion, PDFs, DOCX, emails, Slack exports, etc.
- Extract text + metadata:
- title
- author
- date
- source URL
- permissions / ACLs
- department, project, tags
-
Chunk the text
- Split documents into small overlapping chunks, often ~300–800 tokens.
- Keep structure if possible:
- headings
- paragraphs
- tables
- bullet lists
-
Embed and index
- Convert each chunk into an embedding vector.
- Store vectors in a vector database or search engine:
- pgvector
- Pinecone
- Weaviate
- Milvus
- Elasticsearch/OpenSearch vector search
-
Query-time retrieval
- When a user asks a question:
- embed the query
- retrieve top matching chunks
- filter by permissions so the user only sees allowed docs
- Often use:
- hybrid search (BM25 + vectors)
- reranking for better precision
- When a user asks a question:
-
Answer generation
- Send the question + retrieved chunks to an LLM.
- Instruct it to answer only from provided context.
- Return citations/snippets to make answers trustworthy.
-
Security and privacy
- Enforce document-level and chunk-level access control.
- Log access.
- Encrypt data at rest and in transit.
- Decide whether inference happens:
- in your cloud/VPC
- on-prem
- via a private API with no training on your data
- Redact secrets/PII if needed.
Recommended approach
Best practical option: RAG
Use RAG if:
- docs change often
- you need citations
- you need up-to-date answers
- you need per-user permissions
Fine-tuning is usually not the right first step
Fine-tuning is better for:
- style
- format
- classification
- domain-specific behavior
It is not ideal for storing private company knowledge because:
- hard to update
- harder to cite sources
- can memorize sensitive info
- doesn’t respect doc permissions naturally
Core components
1) Document ingestion pipeline
You need:
- connectors
- text extraction
- OCR for scanned PDFs
- metadata normalization
- deduplication
- versioning
Tools:
- Apache Tika
- Unstructured
- OCR: Tesseract, AWS Textract, Google Document AI
2) Chunking strategy
Good chunking matters a lot.
Common practices:
- split by heading/section first
- keep chunks semantically coherent
- overlap 10–20% between chunks
- store chunk metadata and parent document ID
Avoid:
- splitting mid-table or mid-sentence too often
- huge chunks that dilute retrieval
- tiny chunks that lose context
3) Retrieval
Use:
- vector similarity search
- keyword search
- hybrid retrieval
- reranking
Why hybrid?
- keyword search is great for exact terms, names, IDs
- vector search is great for semantic matches
4) Prompting the LLM
Your system prompt should enforce:
- use only provided context
- say “I don’t know” if not enough info
- cite sources
- don’t reveal unauthorized info
5) Permissions
This is critical.
Before retrieval:
- identify user
- determine allowed documents/chunks
- filter search results by ACLs
Never rely on the model to “not answer” confidentially accessible content by itself.
Minimal implementation stack
A common stack:
- Frontend: React / Next.js
- API: FastAPI / Node.js
- Auth: Okta / Azure AD / Auth0 / SSO
- Document processing: Python + Unstructured + Tika
- Embeddings: OpenAI, Cohere, Voyage, or self-hosted models
- Vector DB: pgvector or Pinecone
- LLM: GPT-4.1 / Claude / Llama 3.1 via private deployment
- Observability: Langfuse / OpenTelemetry
If privacy is paramount:
- self-host embeddings and LLMs
- keep everything inside your VPC
- avoid sending raw docs to third-party APIs unless contractually acceptable
Data model suggestions
Store something like:
document_idchunk_idtextembeddingsourcetitleurlcreated_atupdated_atacl_usersacl_groupsdepartmentversion
At query time, filter by:
- user ID
- groups
- document labels
- project membership
Answering flow example
- User asks: “What’s our PTO policy for contractors?”
- System authenticates user.
- Search only documents user can access.
- Retrieve top 5–20 chunks.
- Rerank them.
- Give the best chunks to the LLM.
- LLM answers:
- concise response
- citations
- notes if policy differs by region or contract type
Important guardrails
- Access control first, then retrieval
- Citations required
- No answer without evidence for policy/legal questions
- Versioning so old policy docs don’t override new ones
- Feedback loop to capture bad answers
- PII and secrets handling
- Audit trail for who asked what
Common failure modes
- bad chunking
- missing metadata
- ignoring permissions
- using only vector search
- no reranker
- stale documents
- hallucinated answers
- poor OCR on scanned docs
- tables not parsed correctly
Quick MVP plan
Week 1
- connect to one source, like PDFs in S3 or Google Drive
- extract text
- chunk and embed
- store in pgvector
- build a simple chat UI
Week 2
- add auth
- add ACL filtering
- add citations
- add reranking
- add feedback buttons
Week 3+
- add more connectors
- improve parsing
- add document versioning
- add monitoring and evals
- add on-prem/private model deployment if needed
If you want the fastest path
Use:
- LlamaIndex or LangChain for orchestration
- pgvector for storage
- OpenSearch for hybrid search
- FastAPI backend
- SSO + ACL filtering
- private deployment of the LLM if required
If you want, I can also give you:
- a reference architecture diagram,
- a step-by-step MVP implementation, or
- a sample code stack in Python for private document Q&A.