Prompt
How do I build a RAG app that connects to my company documents?
Latest observation
Building a RAG app for company documents usually means:
- Ingest documents
- Chunk them
- Create embeddings
- Store them in a vector database
- Retrieve relevant chunks at query time
- Pass them to an LLM to generate an answer with citations
Here’s a practical blueprint.
1) Define the scope and access rules first
Before coding, decide:
- What sources: PDFs, Google Drive, SharePoint, Confluence, Slack, Notion, web pages, ticketing systems, etc.
- Who can see what: document-level and row-level permissions matter a lot.
- Update frequency: hourly, daily, near-real-time.
- Answer style: short answers, summarized answers, citations, policy compliance, etc.
If your company has sensitive docs, permission filtering is not optional. Your retriever must only return docs the user is allowed to access.
2) Choose a basic architecture
A common stack:
- Document connectors: pull from Drive/SharePoint/S3/etc.
- Parsing/OCR: extract text from PDFs, scans, Office docs
- Chunking pipeline: split documents into searchable pieces
- Embedding model: convert chunks to vectors
- Vector database: Pinecone, Weaviate, pgvector, Milvus, OpenSearch, FAISS
- RAG API: search + prompt + LLM response
- UI: chat interface, search, citations, feedback
A simple flow:
Company docs -> ingest -> clean text -> chunk -> embed -> vector DB
User question -> embed -> retrieve top-k chunks -> prompt LLM -> answer
3) Ingest documents
You need a job that regularly fetches documents from your sources.
Examples:
- Google Drive API
- Microsoft Graph API for SharePoint/OneDrive
- Confluence REST API
- S3 bucket sync
- Web crawler for internal docs site
Store metadata for each doc:
- document ID
- title
- source system
- URL
- owner
- department
- timestamps
- ACL / group permissions
This metadata is crucial for filtering and citations.
4) Parse and normalize text
Documents often contain messy content:
- PDF layout issues
- tables
- headers/footers
- OCR errors
- duplicated text
Use tools like:
unstructuredpymupdfpdfplumberpython-docx- OCR: Tesseract, Azure OCR, AWS Textract, Google Document AI
Clean up:
- remove boilerplate
- deduplicate repeated headers/footers
- preserve section headings
- keep tables if they matter
5) Chunk the text intelligently
Don’t embed entire documents. Split into chunks.
Typical starting point:
- Chunk size: 300–800 tokens
- Overlap: 50–150 tokens
Best practices:
- chunk by headings/sections when possible
- keep tables together
- preserve document title and section path in metadata
- avoid splitting on arbitrary character boundaries only
Example metadata per chunk:
{
"doc_id": "hr-policy-2025",
"chunk_id": "hr-policy-2025-004",
"title": "Remote Work Policy",
"section": "Time Tracking",
"source_url": "https://...",
"permissions": ["hr", "managers"],
"updated_at": "2026-01-15"
}
6) Create embeddings
Use an embedding model to convert each chunk into a vector.
Options:
- OpenAI embeddings
- Cohere embeddings
- Voyage embeddings
- SentenceTransformers / local models
For enterprise doc search, strong general-purpose embeddings are usually enough to start.
Important:
- Use the same embedding model for indexing and querying
- Re-embed if the model changes
- Consider multilingual embeddings if docs are multilingual
7) Store in a vector database
You need a database that supports:
- vector similarity search
- metadata filtering
- updates/deletes
- scalability
Common choices:
- Postgres + pgvector: great for simpler deployments
- Pinecone: managed, easy
- Weaviate: flexible, hybrid search
- Milvus: scalable open-source
- OpenSearch/Elasticsearch: useful if combining keyword + vector search
Recommended pattern:
- store vectors
- store chunk text
- store metadata
- store ACL info for filtering
8) Implement retrieval
At query time:
- Embed user question
- Search vector DB for top-k similar chunks
- Apply metadata filters:
- user permissions
- document type
- business unit
- date range
- Optionally re-rank results with a cross-encoder or LLM reranker
- Send the best chunks to the LLM
You’ll usually get better results by combining:
- vector search
- keyword/BM25 search
- reranking
This is called hybrid retrieval and is often better than vector-only.
9) Build the answer prompt
Your prompt should:
- instruct the model to use only retrieved context
- answer concisely
- cite sources
- say when it doesn’t know
Example pattern:
You are a helpful assistant for internal company documents.
Answer the question using only the provided context.
If the context is insufficient, say you don’t know.
Cite sources for each claim.
Context:
[1] ...
[2] ...
Question: ...
Answer:
For citations, include chunk IDs or source URLs and make the model output them inline or in a references section.
10) Handle permissions securely
This is one of the biggest enterprise pitfalls.
You need:
- document ACLs
- user identity propagation
- group membership checks
- metadata filtering in retrieval
A safe pattern:
- index each chunk with permitted users/groups
- at query time, filter retrieval by the current user’s access rights
- never rely only on the LLM to “not mention” restricted info
Do not fetch all docs and filter after generation. Filter before retrieval.
11) Improve quality with reranking and evaluation
RAG quality is often limited by retrieval, not generation.
Add:
- hybrid retrieval
- reranking
- query rewriting
- multi-hop retrieval for complex questions
- context compression if chunks are too large
Evaluate with:
- answer correctness
- citation accuracy
- retrieval recall@k
- latency
- permission leakage tests
Make a small test set of real employee questions and compare outputs.
12) Add update and deletion handling
Company docs change often.
You need:
- document versioning
- re-ingestion on update
- delete chunks when documents are removed
- incremental sync if possible
Good practice:
- use a stable
doc_id - store
content_hash - only re-embed if content changed
- soft-delete old chunks or version them
13) Build the app interface
Typical UI features:
- chat
- source citations
- “open document” links
- feedback buttons
- suggested questions
- document search mode
Also include:
- confidence/“I’m not sure” behavior
- logs for retrieval debugging
- admin tools to monitor failed ingestion jobs
14) Recommended starter stack
If you want something practical and fast:
Simple enterprise-friendly stack
- Backend: Python + FastAPI
- Parsing:
unstructured+ OCR as needed - Embeddings: OpenAI / Cohere / local embedding model
- Vector DB: Postgres + pgvector
- LLM: GPT-style model or equivalent
- Orchestration: Celery / cron / Airflow / Prefect
- Frontend: Next.js / React
If you want a fast prototype
- LangChain or LlamaIndex for orchestration
- pgvector or Pinecone for retrieval
- a simple chat UI
- document connector for just one source first, like Google Drive
15) A minimal implementation outline
Pseudo-flow:
# ingest.py
docs = fetch_documents_from_drive()
for doc in docs:
text = parse_document(doc)
chunks = chunk_text(text)
for chunk in chunks:
vector = embed(chunk.text)
vector_db.upsert(
id=chunk.id,
vector=vector,
metadata={
"doc_id": doc.id,
"title": doc.title,
"url": doc.url,
"permissions": doc.permissions
},
text=chunk.text
)
# query.py
def answer_question(user প্রশ্ন, user_permissions):
qvec = embed(user_question)
chunks = vector_db.search(
vector=qvec,
top_k=5,
filter={"permissions": {"$in": user_permissions}}
)
prompt = build_prompt(chunks, user_question)
return llm.generate(prompt)
16) Common mistakes to avoid
- No access control
- Chunks too large or too small
- Using only vector search
- No citations
- Ignoring OCR/layout quality
- Not evaluating retrieval
- Not handling document updates/deletions
- Letting the LLM answer without enough context
- Using raw docs with sensitive info without governance
17) If you want to start this week
A good 3-step MVP:
- Pick one source: e.g. Google Drive or SharePoint
- Build ingestion + pgvector
- Create a chat UI with citations and permission filtering
Then improve:
- better chunking
- hybrid search
- reranking
- feedback loop
- more connectors
If you want, I can also give you:
- a reference architecture diagram,
- a Python starter project structure, or
- a working FastAPI + pgvector example for company docs.