Prompt
How do I turn multi-page PDFs into searchable text with OCR?
Latest observation
Oct 11, 2026
OpenAI APIWeb search: off
To turn multi-page PDFs into searchable text with OCR, you generally need to:
- Extract each page as an image if the PDF is scanned or image-based.
- Run OCR on those images to detect text.
- Save the result as searchable text or a new PDF with an invisible text layer.
Easiest options
1) Adobe Acrobat
- Open the PDF
- Go to Tools → Scan & OCR
- Choose Recognize Text → In This File
- Save the PDF
This creates a searchable PDF.
2) Free command-line tool: ocrmypdf
This is one of the best options for multi-page PDFs.
Install
- macOS:
brew install ocrmypdf - Ubuntu/Debian:
sudo apt install ocrmypdf - Windows: use Python/pip or install via package managers; Ghostscript and Tesseract are also needed
Run OCR
ocrmypdf input.pdf output_searchable.pdf
Optional:
ocrmypdf --deskew --clean input.pdf output_searchable.pdf
This:
- processes all pages
- adds a text layer
- keeps the original page images
- makes the PDF searchable
If you want plain text instead of searchable PDF
Use Tesseract on each page image
- Convert PDF pages to images:
pdftoppm -png input.pdf page
- OCR each page:
tesseract page-1.png page-1
- Combine text files:
cat page-*.txt > output.txt
Programmatic option in Python
If you want to automate it:
import pytesseract
from pdf2image import convert_from_path
pages = convert_from_path("input.pdf", dpi=300)
all_text = []
for i, page in enumerate(pages, start=1):
text = pytesseract.image_to_string(page)
all_text.append(f"--- Page {i} ---\n{text}")
with open("output.txt", "w", encoding="utf-8") as f:
f.write("\n\n".join(all_text))
You’ll need:
pytesseractpdf2image- Poppler
- Tesseract OCR installed on your system
Tips for better OCR
- Use 300 DPI or higher
- Ensure pages are deskewed
- Improve contrast if scans are faint
- Use the correct language:
tesseract input.png output -l eng
If you want, I can give you:
- a one-line command for your OS,
- a Python script for batch OCR,
- or help you choose the best tool for searchable PDF vs plain text.
Brands