Prompt
How do I extract structured fields from invoices with an AI API?
Latest observation
To extract structured fields from invoices with an AI API, you typically use one of these approaches:
1) Use a document understanding API
Best for production invoice extraction.
These APIs are built to detect layout and pull fields like:
- invoice number
- date
- vendor name
- bill-to / ship-to
- line items
- subtotal, tax, total
- currency
Examples:
- Google Document AI
- AWS Textract
- Azure Form Recognizer / Document Intelligence
- Mindee
- Veryfi
Typical flow
- Upload invoice PDF/image
- Send it to the API
- Get back JSON with extracted fields
- Validate/normalize the values in your app
Example output
{
"invoice_number": "INV-1024",
"invoice_date": "2026-01-12",
"vendor_name": "Acme Supplies Inc.",
"total": 145.67,
"currency": "USD",
"line_items": [
{"description": "Paper", "quantity": 2, "unit_price": 12.50, "amount": 25.00}
]
}
2) Use a general LLM with structured output
Best if invoices vary a lot or you already use an LLM platform.
Workflow
- Convert invoice to text or OCR it first
- Prompt the model to extract fields
- Require JSON output using a schema
- Validate results
Example prompt
Extract the following invoice fields from the text below and return valid JSON matching this schema: invoice_number, invoice_date, vendor_name, bill_to, subtotal, tax, total, currency, line_items[].
Example schema
{
"type": "object",
"properties": {
"invoice_number": {"type": "string"},
"invoice_date": {"type": "string"},
"vendor_name": {"type": "string"},
"total": {"type": "number"}
},
"required": ["invoice_number", "invoice_date", "vendor_name", "total"]
}
This is especially useful when:
- invoices are semi-structured
- you need custom fields
- you want one API for many document types
3) Hybrid: OCR + LLM post-processing
Very common in real systems.
Why?
- OCR handles text extraction from scans/photos
- LLM cleans up and maps fields into your schema
- You can add business rules after extraction
Example pipeline
- OCR invoice image → text
- LLM extracts JSON
- Validate totals, dates, currency
- Flag low-confidence invoices for human review
Recommended architecture
For most applications:
- OCR/document AI for text extraction
- LLM for normalization and schema mapping
- Validation layer to check:
- totals = subtotal + tax
- dates parse correctly
- currency is valid
- required fields are present
- Human review for low-confidence cases
Minimal API example
If you're using an LLM API with structured output, the idea looks like this:
schema = {
"type": "object",
"properties": {
"invoice_number": {"type": "string"},
"invoice_date": {"type": "string"},
"vendor_name": {"type": "string"},
"total": {"type": "number"}
},
"required": ["invoice_number", "invoice_date", "vendor_name", "total"]
}
Then send invoice text along with instructions to return JSON matching the schema.
Tips for better extraction
- Use high-quality scans or PDFs
- Run OCR before the LLM if the source is an image
- Ask for strict JSON
- Include examples in the prompt
- Normalize dates/currency after extraction
- Use confidence scores if your API provides them
- Build fallback rules for missing fields
If you want accuracy on invoices specifically
Choose a document AI tool if:
- you need line items
- invoices are scanned and messy
- you care about consistent extraction at scale
Choose an LLM if:
- you need flexibility
- invoice formats vary widely
- you want custom logic or extra fields
If you want, I can show you:
- a Python example using an invoice OCR/API, or
- a prompt + JSON schema for invoice extraction with an LLM.