Prompt
Need LLM API for structured extraction from PDFs
Latest observation
Oct 9, 2026
OpenAI APIWeb search: off
If you want an LLM API for structured extraction from PDFs, the best choice depends on:
- Whether the PDF is text-based or scanned
- How structured the output must be
- Whether you need just extraction or also OCR, tables, layout, and citations
Good API options
1. OpenAI API
Best if you want:
- Strong extraction from text PDFs
- JSON / schema-based structured output
- Reasoning over document content
Typical flow:
- Convert PDF pages to text or images
- Send page text/images to the model
- Ask for strict JSON output
Good for:
- Invoices, contracts, reports, forms
- Key-value extraction
- Summarization + structure
Limitations:
- You usually need your own PDF parsing/OCR pipeline for best results
2. Google Document AI
Best if you want:
- Native document understanding
- OCR + layout-aware extraction
- Tables, forms, and scanned PDFs
Good for:
- Enterprise document pipelines
- High-volume extraction
- Prebuilt processors for invoices, IDs, receipts, etc.
3. Azure Document Intelligence
Best if you want:
- OCR and structured extraction with form understanding
- Table extraction
- Easy enterprise integration
Good for:
- Business documents
- Scanned PDFs
- Custom models for your document types
4. Anthropic API
Best if you want:
- Strong reasoning on extracted text
- Reliable structured output
- Long-context document processing
Still typically paired with:
- PDF parsing
- OCR for scanned docs
Recommended architecture
For most use cases:
Text PDFs
- Extract text with
pdfplumber,pymupdf, orpdftotext - Chunk by page or section
- Send to LLM
- Enforce JSON schema output
Scanned PDFs
- OCR with Document AI / Azure / Tesseract
- Pass OCR text + layout hints to LLM
- Extract into structured JSON
If you want the simplest LLM-based structured extraction
Use:
- OpenAI API for the extraction step
- PyMuPDF or pdfplumber for PDF text extraction
- OCR only if needed
This is often the most flexible and cost-effective approach.
Example output shape
You can ask the model to return:
{
"document_type": "invoice",
"vendor_name": "ABC Corp",
"invoice_number": "INV-12345",
"invoice_date": "2026-01-01",
"total_amount": 1234.56,
"currency": "USD",
"line_items": [
{
"description": "Consulting services",
"quantity": 10,
"unit_price": 100.0,
"amount": 1000.0
}
]
}
My recommendation
- Need maximum document/OCR accuracy: Google Document AI or Azure Document Intelligence
- Need flexible schema extraction with LLM reasoning: OpenAI API
- Need both: use OCR/document extraction first, then LLM for normalization and final JSON
If you want, I can also give you:
- a Python example using OpenAI for PDF structured extraction, or
- a comparison table of OpenAI vs Google Document AI vs Azure.