Prompt
How do I extract text from scanned PDFs with an OCR API?
Latest observation
Oct 11, 2026
OpenAI APIWeb search: off
To extract text from scanned PDFs with an OCR API, the general flow is:
- Upload the PDF
- Send it to the OCR service
- Wait for processing
- Retrieve the extracted text
- Clean/post-process the result
Typical workflow
1) Choose an OCR API
Common options:
- Google Cloud Vision / Document AI
- AWS Textract
- Azure AI Document Intelligence
- Tesseract if you want local/open-source OCR
For scanned PDFs, use a service that supports document OCR, not just image OCR.
2) Send the PDF to the API
Most OCR APIs accept either:
- a file upload
- a public/private URL
- or a file in cloud storage like S3 / GCS / Azure Blob
For large PDFs, many APIs use asynchronous processing:
- submit the job
- get a job ID
- poll until complete
- download results
3) Parse the response
The OCR response usually contains:
- full extracted text
- page-by-page text
- bounding boxes/coordinates
- confidence scores
- sometimes tables/forms/paragraph structure
Example approach with a generic OCR API
Python example
import requests
api_url = "https://api.example.com/ocr"
api_key = "YOUR_API_KEY"
with open("scanned.pdf", "rb") as f:
files = {"file": f}
headers = {"Authorization": f"Bearer {api_key}"}
response = requests.post(api_url, files=files, headers=headers)
print(response.json())
If the API is async:
import requests
import time
submit_url = "https://api.example.com/ocr/jobs"
status_url = "https://api.example.com/ocr/jobs/{job_id}"
with open("scanned.pdf", "rb") as f:
r = requests.post(submit_url, files={"file": f}, headers={"Authorization": "Bearer YOUR_API_KEY"})
job_id = r.json()["job_id"]
while True:
status = requests.get(status_url.format(job_id=job_id),
headers={"Authorization": "Bearer YOUR_API_KEY"}).json()
if status["state"] == "done":
print(status["text"])
break
elif status["state"] == "failed":
raise RuntimeError(status["error"])
time.sleep(2)
If the PDF is image-based
If the scanned PDF is not already OCR’d, the API will:
- render each page to an image internally, or
- require you to convert pages to images first
Sometimes converting pages yourself can help if the API only accepts images:
pdf2imagepopplerimagemagick
Best practices
- Preprocess scans for better accuracy:
- deskew
- increase contrast
- remove noise
- ensure at least 300 DPI
- Use async jobs for multi-page PDFs
- Handle page order
- Check confidence scores
- Use language settings if available
Common pitfalls
- Uploading a PDF to an API that only accepts images
- Not using async mode for large files
- Poor scan quality
- Trying to OCR text-based PDFs when you could extract text directly without OCR
If you want plain text only
Many APIs return structured JSON. You can usually combine page text like this:
text = "\n".join(page["text"] for page in result["pages"])
If you want, I can show you:
- a real example using Google, AWS Textract, or Azure
- a Python script for OCRing a PDF end-to-end
- how to extract text locally with Tesseract