Prompt
How do I connect an LLM to our internal docs so it answers from our data instead of guessing?
Latest observation
The standard way is to use Retrieval-Augmented Generation (RAG).
What RAG does
Instead of letting the LLM rely only on its pretraining, you:
- Index your internal documents
- Retrieve the most relevant passages for a user question
- Give those passages to the LLM as context
- Tell the LLM to answer only from that context
That makes the model answer from your data instead of guessing.
Typical architecture
1) Ingest documents
Collect sources like:
- PDFs
- Confluence / Notion
- Google Docs
- SharePoint
- Wikis
- Tickets / runbooks / internal FAQs
Extract the text and metadata:
- title
- author
- date
- source URL
- access permissions
- department/team
2) Chunk the content
Split documents into smaller pieces, usually:
- 300–800 tokens each
- with overlap, like 10–20%
This helps retrieval find the exact relevant section instead of an entire long doc.
3) Create embeddings
Convert each chunk into a vector embedding using an embedding model.
Store:
- the chunk text
- its embedding
- metadata
4) Put them in a vector database
Examples:
- Pinecone
- Weaviate
- Milvus
- Qdrant
- FAISS
- pgvector in Postgres
5) At question time
When a user asks something:
- Embed the question
- Find the most similar chunks in the vector DB
- Optionally rerank those results
- Send the top chunks to the LLM in the prompt
6) Prompt the LLM carefully
Use instructions like:
- “Answer only using the provided context.”
- “If the answer is not in the context, say you don’t know.”
- “Cite the source chunks.”
Example flow
User asks:
“What is our vacation policy for contractors?”
System:
- searches internal docs
- retrieves the policy section
- passes it to the LLM
LLM output:
“Contractors receive 10 unpaid vacation days per year. Requests must be approved by the project manager at least 2 weeks in advance. Source: HR Policy v3, section 4.2.”
If no relevant doc is found:
“I couldn’t find this in the provided documents.”
Important design choices
Retrieval quality matters a lot
Good answers depend on:
- good chunking
- strong embeddings
- metadata filters
- reranking
Use access control
If users have different permissions, retrieval must respect them. Don’t let the model see docs the user shouldn’t access.
Add citations
Always return source links or document IDs so users can verify the answer.
Don’t fine-tune first
For internal docs, RAG is usually better than fine-tuning because:
- easier to update
- more accurate for current info
- more transparent
- less expensive to maintain
Common stack
A simple production setup might be:
- Document loader: custom scripts, Unstructured, Apache Tika
- Chunking: LangChain or LlamaIndex
- Embeddings: OpenAI, Cohere, bge, e5
- Vector DB: pgvector, Pinecone, Qdrant
- LLM: GPT-4.1, Claude, Llama, etc.
- Orchestration: LangChain, LlamaIndex, custom service
Best practices
- Keep chunks semantically coherent
- Use metadata filters by department, date, or ACL
- Add reranking for better precision
- Return citations
- Log unanswered queries to improve coverage
- Refresh the index when docs change
- Evaluate with a test set of real internal questions
Minimal prompt pattern
You can use something like:
You are a helpful assistant for internal company knowledge.
Answer only using the context provided below.
If the answer is not in the context, say “I don’t know based on the provided documents.”
Include citations for each claim.
Then provide:
- retrieved context chunks
- the user question
If you want, I can also give you:
- a simple architecture diagram,
- a Python example using RAG, or
- a production checklist for internal doc search.