Prompt

How do I extract text from scanned PDFs with an OCR API?

Artificial Intelligence / Computer Vision1 observationLast seen Oct 11, 2026

Latest observation

Oct 11, 2026OpenAI APIWeb search: off

To extract text from scanned PDFs with an OCR API, the general flow is:

  1. Upload the PDF
  2. Send it to the OCR service
  3. Wait for processing
  4. Retrieve the extracted text
  5. Clean/post-process the result

Typical workflow

1) Choose an OCR API

Common options:

  • Google Cloud Vision / Document AI
  • AWS Textract
  • Azure AI Document Intelligence
  • Tesseract if you want local/open-source OCR

For scanned PDFs, use a service that supports document OCR, not just image OCR.


2) Send the PDF to the API

Most OCR APIs accept either:

  • a file upload
  • a public/private URL
  • or a file in cloud storage like S3 / GCS / Azure Blob

For large PDFs, many APIs use asynchronous processing:

  • submit the job
  • get a job ID
  • poll until complete
  • download results

3) Parse the response

The OCR response usually contains:

  • full extracted text
  • page-by-page text
  • bounding boxes/coordinates
  • confidence scores
  • sometimes tables/forms/paragraph structure

Example approach with a generic OCR API

Python example

import requests

api_url = "https://api.example.com/ocr"
api_key = "YOUR_API_KEY"

with open("scanned.pdf", "rb") as f:
    files = {"file": f}
    headers = {"Authorization": f"Bearer {api_key}"}
    response = requests.post(api_url, files=files, headers=headers)

print(response.json())

If the API is async:

import requests
import time

submit_url = "https://api.example.com/ocr/jobs"
status_url = "https://api.example.com/ocr/jobs/{job_id}"

with open("scanned.pdf", "rb") as f:
    r = requests.post(submit_url, files={"file": f}, headers={"Authorization": "Bearer YOUR_API_KEY"})

job_id = r.json()["job_id"]

while True:
    status = requests.get(status_url.format(job_id=job_id),
                         headers={"Authorization": "Bearer YOUR_API_KEY"}).json()
    if status["state"] == "done":
        print(status["text"])
        break
    elif status["state"] == "failed":
        raise RuntimeError(status["error"])
    time.sleep(2)

If the PDF is image-based

If the scanned PDF is not already OCR’d, the API will:

  • render each page to an image internally, or
  • require you to convert pages to images first

Sometimes converting pages yourself can help if the API only accepts images:

  • pdf2image
  • poppler
  • imagemagick

Best practices

  • Preprocess scans for better accuracy:
    • deskew
    • increase contrast
    • remove noise
    • ensure at least 300 DPI
  • Use async jobs for multi-page PDFs
  • Handle page order
  • Check confidence scores
  • Use language settings if available

Common pitfalls

  • Uploading a PDF to an API that only accepts images
  • Not using async mode for large files
  • Poor scan quality
  • Trying to OCR text-based PDFs when you could extract text directly without OCR

If you want plain text only

Many APIs return structured JSON. You can usually combine page text like this:

text = "\n".join(page["text"] for page in result["pages"])

If you want, I can show you:

  • a real example using Google, AWS Textract, or Azure
  • a Python script for OCRing a PDF end-to-end
  • how to extract text locally with Tesseract

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.