Prompt
How do I process receipts with an OCR API and get structured fields?
Latest observation
To process receipts with an OCR API and turn them into structured fields, you generally want a pipeline like this:
1) Capture the receipt image
Get the receipt as one of:
- JPG / PNG photo
- PDF scan
- multi-page PDF if needed
Best results come from:
- good lighting
- flat, unwrinkled paper
- full receipt in frame
- high resolution
2) Send it to an OCR API
Use an OCR service that supports either:
- plain text extraction, or
- document/receipt parsing with built-in fields
Examples of APIs that can do this:
- Google Cloud Vision / Document AI
- AWS Textract
- Azure Document Intelligence
- Mindee
- Veryfi
- OCR.space
- ABBYY
- Nanonets
If the API has a receipt endpoint, use that instead of generic OCR, because it can often return structured fields directly.
3) Extract structured fields
Typical receipt fields you may want:
- merchant / store name
- date and time
- subtotal
- tax
- tip
- total
- currency
- payment method
- line items:
- item name
- quantity
- unit price
- total price
- receipt number / transaction ID
Example JSON structure
{
"merchant_name": "ABC Grocery",
"date": "2026-10-11",
"time": "14:32",
"currency": "USD",
"subtotal": 18.49,
"tax": 1.48,
"total": 19.97,
"line_items": [
{
"description": "Milk",
"quantity": 1,
"unit_price": 3.99,
"amount": 3.99
},
{
"description": "Bread",
"quantity": 2,
"unit_price": 2.25,
"amount": 4.50
}
]
}
4) Post-process and validate
OCR output is often noisy, so you should:
- normalize dates
- convert currency strings to decimals
- validate totals:
subtotal + tax (+ tip) ≈ total - handle missing fields
- deduplicate repeated text
- fix common OCR mistakes like
Ovs0,Svs5
5) Use confidence scores
Many OCR APIs return confidence values per field or word. You can:
- trust high-confidence fields directly
- flag low-confidence fields for review
- fall back to rules or manual correction
Two common approaches
A) Use a receipt-specific OCR API
Best if you want structured fields quickly.
Example flow:
- Upload receipt image
- API returns JSON with parsed fields
- Store JSON in your database
Pros:
- less custom code
- better for receipts/invoices
- structured output
Cons:
- vendor lock-in
- cost per document
B) Use generic OCR + your own parser
Best if you need flexibility or already have OCR text.
Example flow:
- OCR API returns raw text
- Parse text with regex / heuristics / NLP
- Map to fields
Pros:
- flexible
- cheaper in some cases
Cons:
- more engineering effort
- less accurate for complex layouts
Example pseudo-code
Python example
import requests
url = "https://api.example.com/receipt-ocr"
headers = {
"Authorization": "Bearer YOUR_API_KEY"
}
with open("receipt.jpg", "rb") as f:
files = {"file": f}
response = requests.post(url, headers=headers, files=files)
data = response.json()
print(data["merchant_name"])
print(data["total"])
print(data["line_items"])
If you only get raw OCR text
You can parse it with rules like:
import re
text = """
ABC Grocery
2026-10-11 14:32
Subtotal 18.49
Tax 1.48
Total 19.97
"""
total_match = re.search(r"Total\s+(\d+\.\d{2})", text, re.I)
if total_match:
total = float(total_match.group(1))
print(total)
For line items, you may need more advanced parsing because receipts vary a lot in format.
Best practices
- Prefer receipt/document parsing APIs over plain OCR if you need structured fields.
- Use image preprocessing:
- deskew
- crop
- increase contrast
- denoise
- Keep a human review step for low-confidence receipts.
- Normalize everything into a standard schema.
- Store both raw OCR output and cleaned structured data.
Recommended architecture
- Upload receipt
- Preprocess image
- Call OCR/receipt API
- Parse structured response
- Validate fields
- Store in DB
- Send low-confidence items to review
If you want, I can also give you:
- a Python example using a real OCR API
- a Node.js example
- or a recommended receipt JSON schema for your app.