Prompt
I'm building a RAG system for messy PDFs and slide decks. What pipeline should I use?
Latest observation
For messy PDFs and slide decks, I’d use a layout-aware, hybrid RAG pipeline rather than a simple “extract text → chunk → embed” flow.
Recommended pipeline
1) Ingest and classify documents
First, detect the document type and quality:
- Born-digital PDF: has selectable text
- Scanned PDF: needs OCR
- Slide deck / PPT export: needs slide-level structure
- Tables-heavy / forms / diagrams: require special handling
This lets you route each document through the right extraction path.
2) Extract layout-aware content
Use tools that preserve structure, not just plain text.
For PDFs
- Extract:
- text blocks with coordinates
- headings
- paragraphs
- tables
- captions
- footnotes
- page numbers
- If scanned:
- OCR with bounding boxes
- optionally use document AI for reading order
For slide decks
- Extract:
- title
- subtitle
- speaker notes
- bullets
- text boxes
- charts/tables if possible
- slide images for vision-based fallback
Good extractor types
unstructuredpdfplumber/pymupdffor PDFs- OCR engines like Tesseract, Azure Document Intelligence, AWS Textract, Google Document AI
- Vision-capable parsing for charts/diagrams when needed
3) Normalize into a common document model
Convert everything into a unified schema, for example:
doc_idpage_or_slideblock_idblock_type(title, paragraph, table, bullet, caption, etc.)textbboxsection_pathsource_uri
This is important so downstream chunking and retrieval can reason about structure.
4) Clean and reconstruct reading order
Messy docs often have:
- two-column layouts
- headers/footers repeated on every page
- broken hyphenation
- text split across boxes
- tables interleaved with prose
So apply:
- header/footer removal
- dehyphenation
- reading-order reconstruction
- bullet/numbered list merging
- table normalization
- page/slide artifact filtering
This step improves retrieval quality a lot.
5) Chunk by semantic structure, not fixed tokens
Avoid naive 500-token chunks. Use hierarchical, structure-aware chunking:
Chunking strategy
- chunk by section/subsection when available
- keep paragraphs intact
- keep table rows or table groups together
- keep slide title + bullets as one unit
- use overlaps only when necessary
Recommended approach
- create small atomic chunks for retrieval
- also keep parent chunks for context expansion
- store parent-child links
This helps answer precision queries while still allowing rich context on retrieval.
6) Generate multiple representations
For each chunk, store:
- raw text
- cleaned text
- layout metadata
- summary of the chunk
- keywords/entities
- optional table-to-text rendering
- optional image caption / OCR text for figures
These extra views often improve retrieval for messy docs.
7) Embed with hybrid retrieval in mind
Use more than one retrieval method:
- dense embeddings for semantic search
- BM25 / sparse search for exact terms, names, numbers, codes
- optionally multivector / late-interaction retrieval for long or complex chunks
For messy PDFs, hybrid retrieval is usually much better than dense-only.
8) Index at multiple levels
Index:
- chunk level
- section level
- document level
This allows:
- precise retrieval for specific facts
- broader retrieval for summarization or “what does this deck say about X?”
9) Query-time retrieval and reranking
At query time:
- rewrite/expand the query if needed
- retrieve with hybrid search
- rerank results with a cross-encoder or LLM reranker
- expand to neighboring chunks or parent section
- feed the final context to the generator
This step is key for messy documents because the first retrieval hit is often incomplete.
10) Use citation-aware answer generation
Return answers with:
- source document
- page/slide number
- chunk excerpt
- confidence if possible
This is especially important when documents are noisy.
Best-practice architecture
A strong production setup looks like this:
Ingestion
→ doc classification
→ extraction/OCR/layout parse
→ normalization
→ cleaning/reading order
→ semantic chunking
→ embeddings + sparse index
→ metadata store
Query
→ query rewrite
→ hybrid retrieval
→ reranking
→ context expansion
→ grounded answer generation
→ citations
Special handling for PDFs and slides
PDFs
Use page layout signals:
- column detection
- header/footer removal
- table detection
- figure/caption linking
Slide decks
Treat each slide as a mini-document:
- title = top-level anchor
- bullets = core content
- speaker notes = extra context
- figures/charts = optional OCR + captioning
Slides often work best with slide-level chunks plus bullet-level subchunks.
Practical stack suggestions
Open-source
- Parsing:
unstructured,pymupdf,pdfplumber - OCR: Tesseract, PaddleOCR
- Layout: LayoutParser, Detectron2-based models
- Retrieval: Elasticsearch/OpenSearch + vector DB
- Reranking: cross-encoder reranker or LLM-based reranker
Managed
- Azure Document Intelligence
- AWS Textract
- Google Document AI
If your PDFs are very messy, managed document AI often saves a lot of engineering time.
Simple rule of thumb
If the docs are messy, your success depends more on:
- better extraction
- better chunking
- hybrid retrieval
- reranking
than on the choice of LLM.
My recommended default pipeline
If you want a pragmatic default:
- Detect doc type and OCR need
- Parse with a layout-aware extractor
- Clean reading order and remove boilerplate
- Chunk by section/slide/paragraph
- Store chunk + parent context + metadata
- Index with BM25 + dense embeddings
- Rerank top hits
- Answer with citations and page/slide references
If you want, I can also give you:
- a reference architecture diagram
- a Python implementation outline
- or a tool comparison for
unstructuredvspymupdfvs Document AI.