Prompt

What should I use for RAG on PDFs, docs, and spreadsheets?

Artificial Intelligence / AI Search1 observationLast seen Oct 10, 2026

Latest observation

Oct 10, 2026OpenAI APIWeb search: off

For RAG over PDFs, Word docs, and spreadsheets, a good default is:

Best practical stack

  • Document parsing / extraction:

    • Unstructured (great general-purpose option)
    • Apache Tika (solid for lots of formats)
    • PyMuPDF / pdfplumber for PDFs specifically
    • python-docx for Word docs
    • pandas / openpyxl for spreadsheets
  • Chunking / indexing:

    • Chunk by logical structure when possible:
      • headings, paragraphs, tables, rows
    • For spreadsheets, often treat:
      • each sheet as a section
      • each row as a record
      • optionally create a natural-language summary per row/table
  • Embeddings + vector store:

    • Embeddings: OpenAI text-embedding-3-large / 3-small, or a strong open-source model like bge-large
    • Vector DB: pgvector, Pinecone, Weaviate, Qdrant, or Chroma for simpler setups
  • Retrieval:

    • Use hybrid search if possible:
      • dense embeddings + keyword/BM25
    • Add reranking for better precision:
      • Cohere rerank, bge reranker, etc.

Important note for spreadsheets

Spreadsheets are usually where naive RAG breaks down. Don’t just embed the whole file as one blob. Better options:

  • row-level embeddings
  • table-level summaries
  • metadata like sheet name, row numbers, column headers
  • if queries are analytical, sometimes use a SQL tool or pandas agent instead of pure RAG

If you want the simplest good solution

A strong “works well in practice” setup is:

  • Unstructured for parsing
  • LangChain or LlamaIndex for orchestration
  • OpenAI embeddings or bge
  • Qdrant or pgvector
  • hybrid retrieval + reranker

Quick recommendation by document type

  • PDFs: PyMuPDF + OCR if needed + chunk by section
  • Docs: python-docx / Unstructured
  • Spreadsheets: pandas/openpyxl + row/table chunking + metadata
  • Scanned PDFs/images: OCR via Tesseract, Azure OCR, Google Document AI, or AWS Textract

If you want, I can give you:

  1. a best-in-class architecture,
  2. a cheap/open-source stack, or
  3. a Python implementation example.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.