Prompt
What should I use to search PDFs, emails, and wiki pages with RAG?
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
For PDFs, emails, and wiki pages, the best default approach for RAG is:
Use a hybrid retrieval setup
Combine:
- Vector search for semantic matching
- Keyword / full-text search for exact terms, names, IDs, error codes, dates, etc.
This is usually better than using embeddings alone, because:
- PDFs often contain noisy formatting
- Emails have short, messy, conversational text
- Wiki pages often include exact terminology that keyword search catches well
Practical stack options
If you want the simplest strong option
Use a search engine that supports hybrid retrieval:
- Elasticsearch / OpenSearch
- Azure AI Search
- Weaviate
- Pinecone with hybrid support
- Qdrant plus a separate keyword index
- Postgres + pgvector if scale is modest, with a text search column
If you want a lightweight local setup
- Postgres + pgvector + full-text search
- Or SQLite + vector search for small prototypes, though weaker for production
Recommended architecture
- Extract text from PDFs, emails, and wiki pages
- Chunk the text into smaller passages
- Store each chunk with:
- embedding vector
- raw text
- metadata like source, title, author, date, thread, URL
- Retrieve with:
- semantic similarity
- keyword search
- optional reranking
- Feed top results into the LLM
Important extras
- Use a reranker if you need better answer quality
- Preserve metadata so the model can cite sources and filter by date/source
- For PDFs, use good parsing/OCR if documents are scanned
- For emails, keep message/thread structure and timestamps
- For wikis, store page hierarchy and section titles
My short recommendation
If you’re building this today:
- Best general choice: OpenSearch/Elasticsearch with hybrid search
- Best managed easy option: Azure AI Search
- Best open-source semantic-first option: Qdrant or Weaviate
- Best small-scale option: Postgres + pgvector + full-text search
If you want, I can also recommend the best tool based on your scale:
- local prototype
- startup production
- enterprise knowledge base
- self-hosted open source