Prompt
How do I connect an LLM to our internal docs so it answers from our data only?
Latest observation
To make an LLM answer from your internal docs only, you usually build a RAG system: Retrieval-Augmented Generation.
What that means
Instead of letting the model rely on its pretraining, you:
- Store your docs in a searchable index
- Retrieve the most relevant passages for each question
- Give those passages to the LLM
- Instruct the LLM to answer only from the provided context
Typical architecture
Internal docs → chunking → embeddings → vector database / search index → retrieve top passages → LLM prompt → answer
Step-by-step
1) Collect and clean your docs
Pull in PDFs, wiki pages, Notion, SharePoint, Confluence, Google Docs, etc.
Then:
- remove boilerplate if needed
- preserve titles, headings, and metadata
- keep source links, timestamps, permissions, and document IDs
2) Split docs into chunks
LLMs work better with smaller text sections.
Common approach:
- chunk size: ~300–800 tokens
- overlap: ~50–150 tokens
Keep chunks semantically coherent if possible, e.g. by heading/section.
3) Create embeddings
Turn each chunk into a vector using an embedding model.
Store:
- chunk text
- embedding vector
- metadata: source, title, section, ACLs, updated_at, etc.
4) Put them in a retrieval system
Options:
- Vector DB: Pinecone, Weaviate, Milvus, Qdrant, pgvector
- Hybrid search: combine keyword + vector search
- Enterprise search: Elasticsearch/OpenSearch with vectors
Hybrid search is often best for internal docs because exact terms matter.
5) Retrieve relevant chunks at question time
When a user asks a question:
- embed the question
- search the index
- return top-k relevant chunks
- optionally rerank them with a stronger model
6) Prompt the LLM with strict instructions
Tell the model:
- answer only using the supplied context
- if the answer isn’t in the context, say “I don’t know”
- cite sources
- don’t use outside knowledge
Example instruction:
Use only the provided context to answer the question. If the answer is not contained in the context, say you don’t have enough information. Do not guess.
7) Add citations
Return:
- answer
- source snippets
- document links
- chunk IDs
This builds trust and makes auditing easier.
Important guardrails
Enforce permissions
If docs have access controls, retrieval must respect them.
Don’t let the model see chunks the user shouldn’t access.
Limit the context
Only provide the most relevant chunks. Too much context can confuse the model.
Use “abstain” behavior
If retrieval finds nothing strong enough, the system should respond:
- “I couldn’t find that in the docs”
- and optionally suggest related docs
Log retrieval results
Store:
- query
- retrieved chunks
- final answer
- citations
Useful for debugging and evals.
A simple prompt pattern
System message
You are a helpful assistant for internal documentation. Answer only using the provided context. If the answer is not in the context, say you don't know. Cite sources.
User message
Question: How do I reset my VPN token?
Context: [chunk 1] [chunk 2] [chunk 3]
Expected behavior
- Answer from the chunks
- Quote or paraphrase accurately
- Cite the document section
If you want “data only,” note this caveat
No LLM can be perfectly guaranteed to “know nothing else” in the absolute sense, because its base model still contains prior knowledge. But you can make it behave as if it only knows your docs by:
- using RAG
- strict prompts
- citations
- answer refusal when context is missing
- optionally using a smaller local model or fine-tuned model for stronger control
For high-stakes use cases, add:
- answer verification
- citation checking
- policy rules
- human review for certain queries
Common stack choices
Fast to build
- Loader: LangChain / LlamaIndex
- Embeddings: OpenAI, Cohere, sentence-transformers
- Vector DB: Pinecone / Qdrant / pgvector
- LLM: GPT-style model via API
More enterprise/self-hosted
- Ingestion pipeline: Airflow / Dagster
- Search: Elasticsearch/OpenSearch + vectors
- Vector DB: pgvector/Qdrant/Milvus
- LLM: hosted private model or on-prem
Minimal pseudo-flow
doc -> chunk -> embed -> index
question -> embed -> retrieve top chunks -> prompt LLM with chunks -> answer
Best practices
- chunk by section, not arbitrary length only
- add metadata and source links
- use hybrid retrieval + reranking
- enforce ACLs in retrieval
- return citations
- tune with real user questions
- evaluate on a test set of known Q&A
If you want, I can also give you:
- a reference architecture diagram,
- a Python example using LangChain or LlamaIndex, or
- a production checklist for secure internal-doc QA.