Prompt
What should I use for RAG on PDFs, docs, and spreadsheets?
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
For RAG over PDFs, Word docs, and spreadsheets, a good default is:
Best practical stack
-
Document parsing / extraction:
- Unstructured (great general-purpose option)
- Apache Tika (solid for lots of formats)
- PyMuPDF / pdfplumber for PDFs specifically
- python-docx for Word docs
- pandas / openpyxl for spreadsheets
-
Chunking / indexing:
- Chunk by logical structure when possible:
- headings, paragraphs, tables, rows
- For spreadsheets, often treat:
- each sheet as a section
- each row as a record
- optionally create a natural-language summary per row/table
- Chunk by logical structure when possible:
-
Embeddings + vector store:
- Embeddings: OpenAI text-embedding-3-large / 3-small, or a strong open-source model like bge-large
- Vector DB: pgvector, Pinecone, Weaviate, Qdrant, or Chroma for simpler setups
-
Retrieval:
- Use hybrid search if possible:
- dense embeddings + keyword/BM25
- Add reranking for better precision:
- Cohere rerank, bge reranker, etc.
- Use hybrid search if possible:
Important note for spreadsheets
Spreadsheets are usually where naive RAG breaks down. Don’t just embed the whole file as one blob. Better options:
- row-level embeddings
- table-level summaries
- metadata like sheet name, row numbers, column headers
- if queries are analytical, sometimes use a SQL tool or pandas agent instead of pure RAG
If you want the simplest good solution
A strong “works well in practice” setup is:
- Unstructured for parsing
- LangChain or LlamaIndex for orchestration
- OpenAI embeddings or bge
- Qdrant or pgvector
- hybrid retrieval + reranker
Quick recommendation by document type
- PDFs: PyMuPDF + OCR if needed + chunk by section
- Docs: python-docx / Unstructured
- Spreadsheets: pandas/openpyxl + row/table chunking + metadata
- Scanned PDFs/images: OCR via Tesseract, Azure OCR, Google Document AI, or AWS Textract
If you want, I can give you:
- a best-in-class architecture,
- a cheap/open-source stack, or
- a Python implementation example.