Prompt
How do I build a private Q&A app over our internal knowledge base?
Latest observation
A good private Q&A app over an internal knowledge base is usually built as a RAG system: Retrieval-Augmented Generation. The idea is:
- Index your internal docs
- Retrieve the most relevant passages for a question
- Send only those passages to the model
- Have the model answer with citations
That keeps answers grounded in your data and avoids training a model on sensitive content.
Recommended architecture
1) Data ingestion
Collect documents from:
- Wikis: Confluence, Notion, SharePoint
- File storage: Google Drive, OneDrive, S3, network drives
- Ticketing systems: Jira, Zendesk
- Internal manuals, PDFs, docs, spreadsheets, HTML
Pipeline:
- Extract text
- Clean and normalize
- Split into chunks
- Attach metadata:
- title
- source URL/path
- author
- department
- created/updated time
- access permissions
2) Indexing
Create two indexes:
- Keyword index for exact matches
- Vector index for semantic search
Store:
- Chunk text
- Embeddings
- Metadata and ACLs
Common choices:
- Vector DB: Pinecone, Weaviate, Milvus, pgvector
- Search: Elasticsearch/OpenSearch
- Embeddings: OpenAI embeddings or open-source alternatives
3) Access control
This is critical for “private.”
Enforce permissions at retrieval time:
- Only retrieve chunks the user is allowed to see
- Use document-level or chunk-level ACLs
- Sync with your identity provider:
- Okta
- Azure AD / Entra ID
- Google Workspace
- LDAP
Rule of thumb: never rely only on the UI for security. The retrieval layer must filter by user permissions.
4) Query pipeline
When a user asks a question:
- Authenticate the user
- Determine their permissions/groups
- Rewrite or expand the query if needed
- Retrieve top relevant chunks
- Rerank results
- Send the best chunks to the LLM
- Generate answer with citations and confidence cues
Good extras:
- Query routing: detect whether the query needs HR docs, engineering docs, policy docs, etc.
- Reranker: improves relevance a lot
- Follow-up memory: store conversation context, but still retrieve fresh evidence
5) Answer generation
Prompt the model to:
- Answer only from the supplied sources
- Say “I don’t know” if sources don’t support an answer
- Cite sources inline
- Avoid hallucinations
Example behavior:
- “According to the PTO policy v3.2, employees accrue 1.5 days/month…”
- If no evidence: “I couldn’t find a policy document confirming that.”
6) User experience
Add:
- Search bar + chat interface
- Source citations with links
- Filters by department/date/type
- Suggested questions
- Feedback buttons:
- helpful/not helpful
- correct/incorrect
- “Open source document” action
7) Admin and observability
You’ll want:
- Audit logs: who asked what, what sources were retrieved, what was answered
- Monitoring: latency, retrieval hit rate, hallucination reports
- Content freshness: re-index on document updates
- Evaluation suite: known Q&A pairs to test accuracy
Suggested technology stack
Simple startup stack
- Backend: Python + FastAPI
- RAG orchestration: LlamaIndex or LangChain
- Vector store: pgvector or Pinecone
- Search: OpenSearch
- Auth: Okta/Entra ID SSO
- LLM: OpenAI or Azure OpenAI
- Frontend: React/Next.js
If you need full open-source / self-hosted
- Backend: FastAPI
- Embeddings: sentence-transformers
- Vector DB: Milvus or pgvector
- LLM: Llama 3 / Mistral via vLLM or Ollama
- Search: OpenSearch
- Auth: SSO + internal RBAC
- Deployment: Kubernetes
Security best practices
- Encrypt data at rest and in transit
- Redact secrets/PII where appropriate
- Separate environments: dev/staging/prod
- Keep prompts and logs free of sensitive raw text when possible
- Restrict model access to only retrieved snippets, not full corpora
- Add rate limiting and abuse detection
- Consider data retention rules
Common mistakes
- Ignoring document permissions
- Chunking too large or too small
- Using only vector search without keyword search
- No reranking
- No citations
- Letting the model answer from memory instead of sources
- Not evaluating accuracy before rollout
Practical implementation plan
Phase 1: MVP
- Ingest 1–2 doc sources
- Basic semantic search
- Chat interface
- Citations
- SSO auth
- ACL filtering
Phase 2: Quality
- Hybrid search + reranking
- Better chunking
- Feedback loop
- Eval set and dashboards
- Conversation memory
Phase 3: Scale
- More sources
- Real-time sync
- Analytics
- Multi-tenant support
- Fine-grained permissions
- Admin tools
Minimal request flow
User question
→ authenticate user
→ fetch user groups/roles
→ retrieve permitted chunks
→ rerank chunks
→ build prompt with sources
→ LLM answers with citations
→ log audit event
Example prompt pattern
You are answering questions using only the provided internal documents.
If the documents do not contain the answer, say you don’t know.
Cite the document titles and passages used.
Do not reveal any information outside the user’s permissions.
If you want, I can also give you:
- a reference architecture diagram,
- a step-by-step implementation plan, or
- a starter code example in Python/FastAPI.