Prompt

How do I handle OCR and messy formatting in RAG for scanned PDFs?

Artificial Intelligence / AI Search1 observationLast seen Oct 10, 2026

Latest observation

Oct 10, 2026OpenAI APIWeb search: off

Handling OCR and messy formatting in RAG for scanned PDFs is mostly about making the text reliable before retrieval, and preserving layout signals that help reconstruct meaning. A good pipeline usually has four layers:


1) Extract text with OCR, but keep structure if possible

For scanned PDFs, plain text extraction is often useless. Use OCR tools that can return:

  • text
  • page numbers
  • bounding boxes
  • reading order
  • confidence scores

Good options:

  • Tesseract: free, decent baseline
  • PaddleOCR: often better on messy scans
  • AWS Textract / Azure Document Intelligence / Google Document AI: best if you can use managed services
  • OCRmyPDF: useful to add a searchable text layer to scanned PDFs

If possible, extract:

  • blocks/paragraphs
  • tables
  • headers/footers
  • figure captions
  • coordinates for each token or line

This helps later when cleaning and chunking.


2) Clean OCR output aggressively

OCR text usually contains:

  • broken line wraps
  • hyphenated words split across lines
  • repeated headers/footers
  • page numbers
  • artifacts like l, I, 0, O confusion
  • duplicated text from multi-column layouts

Typical cleanup steps:

  • Remove headers/footers that repeat across pages
  • Join hyphenated line breaks
  • Reflow paragraphs by merging lines within the same block
  • Fix common OCR errors with domain-specific rules or dictionaries
  • Deduplicate repeated content
  • Normalize whitespace and punctuation
  • Preserve structure markers like section titles, lists, table boundaries

A simple example of useful cleanup:

  • "inter-\nnational" → "international"
  • "1. Introduction\nThis paper..." → keep the heading separate from paragraph text
  • repeated footer "Confidential - Page 3" → remove

3) Chunk by structure, not just fixed length

With messy OCR, fixed-size chunks often split meaning badly. Better chunking strategies:

Prefer structure-aware chunking

Chunk by:

  • page
  • section headings
  • paragraphs
  • table blocks
  • list items

Use overlap carefully

Add modest overlap between chunks, but avoid too much duplication from OCR artifacts.

Keep metadata

Attach:

  • document name
  • page number
  • section heading
  • coordinates or block IDs
  • OCR confidence

That metadata is extremely helpful at retrieval and answer time.


4) Improve retrieval with OCR-aware techniques

OCR text is noisy, so semantic embeddings alone may miss matches. Use hybrid retrieval:

  • BM25 / keyword search for exact terms, names, numbers
  • dense embeddings for semantic similarity

This is especially important for:

  • product codes
  • legal references
  • dates
  • citations
  • numbers and acronyms

Also consider:

  • spell correction before indexing
  • query expansion for common OCR variants
  • lowercasing and normalization
  • indexing both raw OCR text and cleaned text

Practical pipeline recommendation

Ingestion

  1. Convert PDF pages to images if scanned
  2. Run OCR with layout extraction
  3. Detect and remove repeated headers/footers
  4. Reconstruct reading order
  5. Clean OCR text
  6. Chunk by structure
  7. Store chunks with metadata
  8. Index using hybrid retrieval

Retrieval

  1. Normalize user query
  2. Run hybrid search
  3. Rerank top results with a cross-encoder or LLM reranker
  4. Feed best chunks to the generation model
  5. Cite page/section metadata in answers

Common tricks that help a lot

1) Keep both raw and cleaned text

Store:

  • raw OCR text
  • cleaned canonical text

Raw text helps when the model needs to see exact wording; cleaned text helps retrieval.

2) Use confidence scores

If OCR confidence is low:

  • downweight that chunk in retrieval
  • send it through more aggressive cleaning
  • prioritize nearby chunks for context

3) Handle tables separately

Tables are a major OCR failure point. Extract them as:

  • markdown tables
  • row/column records
  • key-value pairs

If table extraction is unreliable, consider treating the whole table as an image and using multimodal/OCR-specific processing.

4) Preserve page and location metadata for citations

This lets the model answer:

  • “According to page 12...”
  • “In section 4.3...”

It also improves trust and debuggability.

5) Use reranking

OCR noise creates many near-matches. A reranker can filter false positives better than embeddings alone.


When formatting is really bad

If the PDF is extremely messy:

  • use OCR page images plus layout detection
  • extract text in blocks instead of lines
  • run a document parser like:
    • Docling
    • Unstructured
    • LayoutParser
    • cloud document AI tools

If reading order is broken by columns, diagrams, or marginal notes, a layout-aware parser is often necessary.


Example approach in practice

A solid stack might be:

  • OCR: PaddleOCR or Textract
  • Cleaning: custom Python rules
  • Chunking: heading/paragraph-aware
  • Indexing: hybrid BM25 + embeddings
  • Reranking: cross-encoder or LLM reranker
  • Answering: LLM with citations

Rule of thumb

For scanned PDFs, RAG quality depends more on document preprocessing quality than on the LLM. If OCR is noisy, invest in:

  1. better OCR
  2. better layout reconstruction
  3. better chunking
  4. hybrid retrieval

If you want, I can give you:

  • a Python preprocessing pipeline
  • a LangChain/LlamaIndex example
  • or a checklist for cleaning OCR’d PDFs before indexing

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.