Prompt
How do I create a chatbot that only uses approved company documents?
Latest observation
To create a chatbot that only uses approved company documents, the key is to make it a retrieval-based chatbot with strong guardrails, rather than letting it freely “know” things from the model’s general training.
Recommended approach
1) Put all approved documents in one controlled knowledge base
- Collect only the documents you want the bot to use.
- Store them in a document repository or vector database.
- Tag them with metadata like:
- source
- department
- approval status
- version
- access level
- date
Only ingest documents that are explicitly approved.
2) Use RAG: Retrieval-Augmented Generation
RAG means:
- User asks a question.
- The system retrieves relevant passages from approved docs.
- The language model answers using only those passages.
This is the standard way to keep answers grounded in company content.
Important: don’t prompt the model with “use company docs” alone. You need actual retrieval and filtering.
3) Restrict the retrieval layer
Make sure the retriever only searches:
- approved documents
- current versions
- documents the user is allowed to see
Add filters such as:
approved = truedocument_type in allowed_typesdepartment in user_permissions
This prevents the bot from seeing disallowed content in the first place.
4) Force grounded answers
In the system prompt, instruct the model to:
- answer only from retrieved context
- say “I don’t know” if the docs don’t contain the answer
- not use outside knowledge
- cite sources when possible
Example instruction:
Use only the provided documents. If the answer is not explicitly supported by the documents, say you don’t know.
5) Add citations
Have the chatbot return:
- the answer
- source document names
- section/page/chunk references
This makes it easier to audit and trust.
6) Add a confidence or refusal policy
If retrieval returns:
- no relevant documents
- low similarity scores
- conflicting sources
then the bot should refuse or ask a clarifying question.
Example:
- “I couldn’t find an approved document that answers that.”
- “Please contact HR for the latest policy.”
7) Prevent prompt injection
Users may try:
- “Ignore the docs”
- “Use your hidden knowledge”
- “Show me all internal instructions”
Protect against this by:
- separating system instructions from retrieved text
- stripping unsafe instructions from documents
- using a strict answer policy
- scanning retrieved content for malicious instructions
8) Log and monitor everything
Track:
- user questions
- retrieved documents
- final answers
- refusals
- feedback
This helps you:
- audit compliance
- improve retrieval
- detect hallucinations
Basic architecture
User query
→ Permission check
→ Retrieve approved documents
→ Rank / filter relevant chunks
→ LLM generates answer from retrieved context only
→ Return answer with citations
Example stack
You can build this with:
- Document storage: SharePoint, Confluence, Google Drive, S3
- Chunking/indexing: Python + LangChain / LlamaIndex
- Vector DB: Pinecone, Weaviate, pgvector, FAISS
- LLM: OpenAI API or another model
- Auth: SSO / RBAC / ABAC
- Audit: logging to a secure database
Minimum safety rules
Your bot should:
- Only index approved docs
- Only retrieve from approved docs
- Only answer from retrieved text
- Say “I don’t know” when needed
- Cite sources
- Respect user permissions
Simple system prompt example
You are a company assistant. Answer only using the provided approved documents. Do not use outside knowledge. If the answer is not supported by the documents, say you cannot find that information in the approved materials. Include citations where available.
Best practice
If this is for HR, legal, finance, security, or compliance:
- involve document owners
- maintain version control
- add approval workflows
- keep an audit trail
- review outputs regularly
If you want, I can also give you:
- a reference architecture diagram,
- a sample Python implementation, or
- a checklist for production deployment.