Prompt
How do I batch OCR multi-page PDFs through an API?
Latest observation
Oct 11, 2026
OpenAI APIWeb search: off
To batch OCR multi-page PDFs through an API, the usual pattern is:
- Upload the PDF
- Create an OCR job
- Poll job status or receive a webhook
- Download the extracted text / searchable PDF / JSON output
- Repeat in bulk with rate limiting and retries
Common workflow
1) Upload files
If the API requires pre-uploading documents, send each PDF first.
Example:
curl -X POST https://api.example.com/v1/files \
-H "Authorization: Bearer YOUR_API_KEY" \
-F "file=@document.pdf"
Response might include a file_id:
{ "file_id": "file_123" }
2) Start OCR jobs in batch
Submit one job per PDF, or a batch request if supported.
Example:
curl -X POST https://api.example.com/v1/ocr/jobs \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"file_id": "file_123",
"output_format": "json",
"language": "eng",
"pages": "all"
}'
Response:
{ "job_id": "job_456", "status": "queued" }
3) Check job status
Poll until complete:
curl -X GET https://api.example.com/v1/ocr/jobs/job_456 \
-H "Authorization: Bearer YOUR_API_KEY"
Possible response:
{
"job_id": "job_456",
"status": "completed",
"result_url": "https://..."
}
4) Download results
Depending on the API, results may be:
- plain text
- JSON with page/line/word coordinates
- searchable PDF
- HOCR / ALTO XML
Example:
curl -L "https://api.example.com/results/job_456" -o output.json
Batch processing pattern
If you have many PDFs, use:
- async job submission
- concurrency limits
- retry with exponential backoff
- webhooks if available instead of polling
Example pseudo-code
import time
import requests
from concurrent.futures import ThreadPoolExecutor
API_KEY = "YOUR_API_KEY"
HEADERS = {"Authorization": f"Bearer {API_KEY}"}
def submit_and_wait(path):
with open(path, "rb") as f:
r = requests.post(
"https://api.example.com/v1/files",
headers=HEADERS,
files={"file": f}
)
file_id = r.json()["file_id"]
r = requests.post(
"https://api.example.com/v1/ocr/jobs",
headers=HEADERS,
json={"file_id": file_id, "output_format": "json", "pages": "all"}
)
job_id = r.json()["job_id"]
while True:
s = requests.get(f"https://api.example.com/v1/ocr/jobs/{job_id}", headers=HEADERS).json()
if s["status"] in ["completed", "failed"]:
return s
time.sleep(2)
pdfs = ["a.pdf", "b.pdf", "c.pdf"]
with ThreadPoolExecutor(max_workers=4) as ex:
results = list(ex.map(submit_and_wait, pdfs))
Things to watch for
- File size limits: large PDFs may need chunking or preprocessing.
- Page limits: some APIs limit pages per job.
- Timeouts: OCR can take time; use async jobs.
- Rate limits: batch with limited concurrency.
- Output format: choose JSON if you need structured extraction; searchable PDF if you need document archives.
- Language/model options: specify language(s) for better accuracy.
- Rotation/de-skew: some APIs have image cleanup options.
If you’re using a specific OCR API
The exact endpoints differ. If you tell me which provider you’re using, I can give you:
- the exact request format
- a working Python/Node/cURL example
- a batching script for multi-page PDFs