Prompt
How do I extract key-value pairs and tables from invoices with OCR?
Latest observation
To extract key-value pairs and tables from invoices using OCR, the usual pipeline is:
- Get the text from the invoice image/PDF
- Detect layout structure (headers, rows, columns, fields)
- Map text into fields like invoice number, date, total
- Parse line-item tables into structured rows
1) OCR the document
Use an OCR engine to convert the invoice into text, ideally with bounding boxes for each word/line.
Common options:
- Tesseract: open source, good for basic OCR
- Google Document AI / Vision OCR
- Azure Form Recognizer
- AWS Textract
- PaddleOCR
- EasyOCR
For invoices, managed APIs often work better because they also return:
- text
- coordinates
- confidence
- tables/form structures
2) Extract key-value pairs
Invoices usually have labels like:
- Invoice Number: 12345
- Date: 2026-10-11
- Total: $1,234.56
Typical ways to extract them:
A. Rule-based matching
If the invoice formats are fairly consistent:
- search for keywords like
Invoice No,Invoice #,Date,Total - grab the nearby text to the right or below
Useful when:
- you have a few vendor templates
- layout is stable
B. Spatial pairing
Using OCR bounding boxes:
- find label words
- locate the closest value horizontally or vertically
- pair them based on position
Example:
- label at
(x1, y1) - nearest text block to the right on same line = value
C. ML/document AI extraction
Use invoice-specific extraction models that identify fields directly:
- invoice number
- vendor name
- subtotal
- tax
- total
- due date
This is better for:
- many vendors
- varied layouts
- noisy scans
3) Extract tables
Invoice line items usually appear in tables with columns like:
- Description
- Quantity
- Unit Price
- Amount
Approach options
A. Use a table extraction tool
Tools like:
- AWS Textract
- Azure Form Recognizer
- Google Document AI can detect tables and rows automatically.
B. Use OCR + layout analysis
If you only have OCR text and coordinates:
- Detect column boundaries
- Group words by row based on y-coordinate
- Assign each word to a column by x-coordinate
- Build rows into records
C. Use PDF table extractors
If invoices are PDFs with embedded text, try:
- Camelot
- Tabula
- pdfplumber These work best on digital PDFs, not scanned images.
4) Recommended practical pipeline
A strong real-world pipeline looks like:
- If PDF has text layer: parse text directly first
- If scanned/image: OCR it
- Run layout/table detection
- Extract key-value fields
- Extract line-item tables
- Normalize outputs into JSON
Example JSON:
{
"invoice_number": "INV-10293",
"invoice_date": "2026-10-11",
"vendor_name": "ABC Supplies",
"total": 1234.56,
"line_items": [
{
"description": "Paper A4",
"quantity": 10,
"unit_price": 5.0,
"amount": 50.0
}
]
}
5) Tips for better accuracy
- Preprocess images: deskew, denoise, increase contrast
- Use high-resolution scans
- Crop headers/footers if they confuse OCR
- Standardize currency/date formats
- Validate totals: subtotal + tax = total
- Use confidence thresholds and manual review for low-confidence fields
6) If you want code
A simple Python approach could be:
- OCR with PaddleOCR or Tesseract
- Parse positions
- Apply heuristics for labels/values
- Use Camelot or Tabula for PDFs with tables
If you want, I can give you:
- a Python example using Tesseract/PaddleOCR,
- a cloud API approach with Textract/Azure/Google, or
- a table + key-value extraction pipeline with sample code.