Prompt
How do I build a chatbot that searches our knowledge base before answering?
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
To build a chatbot that searches your knowledge base before answering, use a retrieval-augmented generation (RAG) setup:
Basic flow
- User asks a question
- Search your knowledge base for relevant passages
- Send those passages + the question to the model
- Model answers using the retrieved content
Core components
- Knowledge source: docs, PDFs, wiki pages, tickets, etc.
- Chunking: split docs into small overlapping sections
- Embeddings: turn chunks into vectors
- Vector database / search index: store and retrieve chunks
- LLM: generates the final answer from retrieved context
Typical architecture
- Ingest documents
- Clean and chunk text
- Create embeddings for each chunk
- Store embeddings in a vector DB
- On each question:
- embed the question
- retrieve top-k similar chunks
- optionally rerank results
- prompt the LLM with the question + retrieved text
Prompt pattern
Use a system prompt like:
- “Answer only using the provided context.”
- “If the context doesn’t contain the answer, say you don’t know.”
- “Cite sources when possible.”
Example structure:
Context:
[retrieved passages]
Question:
[user question]
Instructions:
Answer based only on the context. If missing, say you don’t know.
Recommended implementation steps
- Start with one data source
- Use a vector DB like:
- Pinecone
- Weaviate
- Qdrant
- pgvector
- FAISS for local prototyping
- Add metadata:
- title
- URL
- updated_at
- access permissions
- Return citations so users can verify answers
Important design tips
- Keep chunks small enough for precise retrieval, but not too small to lose context
- Use overlap between chunks
- Add hybrid search if possible:
- keyword search + vector search
- Add permission filtering so users only see docs they can access
- Log queries and retrieval quality for tuning
- Have a fallback:
- “I couldn’t find that in the knowledge base”
Minimal pseudo-code
question = user_input
query_vec = embed(question)
docs = vector_db.search(query_vec, top_k=5)
context = "\n\n".join(doc.text for doc in docs)
prompt = f"""
Answer only using the context below.
Context:
{context}
Question:
{question}
"""
answer = llm.generate(prompt)
If you want faster results
Use a framework such as:
- LlamaIndex
- LangChain
- Haystack
They handle ingestion, retrieval, and prompting scaffolding.
Best practice for production
- Evaluate retrieval quality separately from answer quality
- Cache frequent queries
- Monitor hallucinations
- Refresh embeddings when docs change
If you want, I can give you:
- a simple Python example,
- a production architecture diagram, or
- a step-by-step guide using LangChain or LlamaIndex.