Prompt
How do I use an OCR API for invoice data extraction?
Latest observation
Oct 11, 2026
OpenAI APIWeb search: off
To use an OCR API for invoice data extraction, you generally follow this workflow:
1. Pick an OCR/invoice extraction API
Choose a service that supports:
- Text OCR: extracts raw text from the invoice
- Document parsing / field extraction: identifies structured fields like invoice number, date, vendor, total, tax, line items
- Tables/line items: important for invoices
Common options include:
- Google Document AI
- AWS Textract
- Azure Document Intelligence
- Mindee
- Veryfi
- ABBYY
- Open-source OCR + your own parsing
2. Send the invoice file
Most APIs accept:
- PNG/JPG/TIFF
- Sometimes multi-page documents
Typical upload methods:
multipart/form-data- Base64-encoded file
- Cloud storage URL
3. Receive extracted text or structured JSON
Depending on the API, the response may include:
- Raw OCR text
- Bounding boxes for words/lines
- Prebuilt invoice fields
- Confidence scores
- Line items in table form
4. Map fields to your schema
Usually you want to normalize extracted data into your own structure, for example:
invoice_numberinvoice_datevendor_namevendor_addresssubtotaltaxtotalcurrencyline_items[]
5. Validate the output
OCR is imperfect, so you should check:
- Missing fields
- Confidence thresholds
- Currency/date formats
- Totals matching subtotal + tax
- Duplicate invoices
Example: generic OCR API request
Here’s a simple example using curl with a file upload:
curl -X POST "https://api.example.com/v1/invoices/extract" \
-H "Authorization: Bearer YOUR_API_KEY" \
-F "file=@invoice.pdf"
Example JSON response:
{
"invoice_number": "INV-10021",
"invoice_date": "2026-10-01",
"vendor_name": "Acme Supplies Ltd.",
"subtotal": 120.00,
"tax": 12.00,
"total": 132.00,
"currency": "USD",
"line_items": [
{
"description": "Printer paper",
"quantity": 4,
"unit_price": 20.00,
"amount": 80.00
},
{
"description": "Ink cartridges",
"quantity": 2,
"unit_price": 20.00,
"amount": 40.00
}
],
"confidence": 0.98
}
Example in Python
import requests
url = "https://api.example.com/v1/invoices/extract"
headers = {
"Authorization": "Bearer YOUR_API_KEY"
}
with open("invoice.pdf", "rb") as f:
files = {"file": f}
response = requests.post(url, headers=headers, files=files)
data = response.json()
print(data)
Best practices
- Use an API with invoice-specific extraction, not just plain OCR, if you need structured fields.
- Preprocess images for better results:
- high resolution
- correct orientation
- clear scans
- remove shadows and blur
- Keep a fallback for manual review when confidence is low.
- Log OCR outputs for debugging and model tuning.
- Use idempotency or deduplication if invoices can be re-uploaded.
If you want to build a pipeline
A typical invoice extraction pipeline looks like:
- Upload invoice
- OCR + field extraction
- Normalize fields
- Validate totals/date/vendor
- Store results in database
- Send low-confidence cases to human review
If you want, I can also show you:
- how to do this with Google Document AI, AWS Textract, or Azure
- a Node.js example
- a workflow for line-item extraction and validation