Prompt
How do I use an OCR API on multi-page documents?
Latest observation
Using an OCR API on multi-page documents is usually straightforward, but the exact method depends on whether your input is:
- a PDF with many pages
- a TIFF / image stack
- a folder of separate page images
- or a document that needs to be split first
Common approaches
1) Upload the whole multi-page document
Many OCR APIs accept a multi-page PDF or TIFF directly.
You send the file once, and the API processes all pages, often returning:
- text for each page
- page numbers
- coordinates/bounding boxes
- confidence scores
Typical flow:
- Read the document file
- Send it to the OCR endpoint
- Wait for the job to finish
- Retrieve results page by page
This is the easiest option if the API supports it.
2) Process each page separately
If the OCR API only accepts single images, you need to split the document into pages first.
Workflow:
- Convert PDF/TIFF into individual images
- Send each page image to the OCR API
- Collect and combine the returned text in page order
This is common for APIs that are image-only.
3) Use asynchronous/batch processing for large docs
For long documents, OCR is often handled as a background job rather than a synchronous request.
Typical pattern:
- Upload document
- Receive a job ID
- Poll for status or receive a webhook callback
- Download results when complete
This avoids request timeouts for large PDFs.
What to look for in the API docs
Check whether the OCR API supports:
- PDF/TIFF input
- multi-page documents
- async jobs
- page indexing
- document splitting
- result formats like JSON, searchable PDF, or plain text
Also check limits like:
- max file size
- max number of pages
- supported formats
- rate limits
Example pseudocode
If the API supports multi-page PDF upload
response = ocr_api.upload_document("scan.pdf")
job_id = response["job_id"]
result = ocr_api.wait_for_result(job_id)
for page in result["pages"]:
print(page["page_number"], page["text"])
If you must OCR page images individually
pages = split_pdf_into_images("scan.pdf")
all_text = []
for i, page_img in enumerate(pages, start=1):
result = ocr_api.ocr_image(page_img)
all_text.append(f"--- Page {i} ---\n{result['text']}")
document_text = "\n".join(all_text)
print(document_text)
Best practices
- Preserve page order when combining results
- Keep page metadata so you can map text back to the source
- Use async processing for large documents
- Preprocess images if needed:
- rotate skewed pages
- increase contrast
- remove noise
- convert to grayscale
- Validate output for tables, headers, and footers
If you want searchable output
Some OCR APIs can return:
- plain text
- structured JSON
- searchable PDF
- hOCR / ALTO XML
If your goal is archiving or search, searchable PDF or structured JSON is often better than plain text.
If you tell me which OCR API you’re using (Google Vision, Azure OCR, AWS Textract, Tesseract wrapper API, etc.) and your document type (PDF, TIFF, images), I can show you the exact request format and code example.