Prompt
How do I parse bank statements with an OCR API?
Latest observation
A good way to parse bank statements with an OCR API is to treat it as a document extraction pipeline, not just “OCR text in, transactions out.”
1) Decide what you need to extract
Typical fields from a bank statement:
- Bank name
- Account holder name
- Account number / IBAN (if allowed)
- Statement period
- Opening balance
- Closing balance
- Transaction list:
- date
- description
- debit / credit
- amount
- running balance
2) Use OCR + layout extraction if possible
Plain OCR gives you raw text, but bank statements often rely on layout and tables. Prefer an OCR API that supports:
- Text extraction
- Table detection
- Bounding boxes / coordinates
- Key-value extraction
- PDF support for scanned and digital statements
- Multi-page handling
Examples of OCR/document APIs:
- Google Document AI
- AWS Textract
- Azure Form Recognizer / Document Intelligence
- OCR.space
- Tesseract plus custom parsing
3) Typical workflow
Step A: Upload the statement
Send the PDF or image to the OCR API.
Step B: Get structured output
Ask for:
- full text
- line items
- tables
- page coordinates
Step C: Parse into sections
Bank statements usually have:
- header block
- summary block
- transaction table
Use rules or post-processing to detect those sections.
Step D: Normalize transaction rows
Convert dates, amounts, and signs into a consistent format:
01/02/2026→ ISO2026-02-011,234.56→ numeric1234.56- debit/credit columns → signed amount
Step E: Validate
Check for:
- opening balance + transactions = closing balance
- date ranges
- duplicate rows
- missing amounts
4) Parsing strategies
Strategy 1: Rule-based parsing
Best when statements come from one bank or a small set of banks.
You can:
- detect known labels
- split lines with regex
- identify transaction rows by date pattern
Example transaction row regex:
r'(\d{2}/\d{2}/\d{4})\s+(.*?)\s+(-?\d[\d,]*\.\d{2})'
Strategy 2: Table extraction
If the OCR API returns table rows, map columns directly:
- Date
- Description
- Debit
- Credit
- Balance
This is usually more reliable than parsing raw text.
Strategy 3: Hybrid parsing
Use OCR text for the header and summary, and table extraction for transactions. This is often the best approach.
5) Example pipeline in Python
Here’s a simple conceptual example:
import requests
def ocr_bank_statement(file_path):
url = "https://api.example-ocr.com/extract"
with open(file_path, "rb") as f:
files = {"file": f}
response = requests.post(url, files=files, headers={
"Authorization": "Bearer YOUR_API_KEY"
})
return response.json()
def parse_transactions(ocr_result):
transactions = []
for row in ocr_result.get("tables", [])[0].get("rows", []):
# Example structure only
date = row[0]
description = row[1]
debit = row[2]
credit = row[3]
balance = row[4]
transactions.append({
"date": date,
"description": description,
"debit": debit,
"credit": credit,
"balance": balance
})
return transactions
result = ocr_bank_statement("statement.pdf")
transactions = parse_transactions(result)
print(transactions)
6) Handling messy OCR output
Bank statements often have issues like:
- split lines
- repeated headers on every page
- multi-line descriptions
- misread characters (
Ovs0,Ivs1) - currency symbols removed
- negative values in parentheses
You’ll likely need cleanup logic:
- merge lines that belong to one transaction
- remove page headers/footers
- standardize whitespace
- use confidence scores to flag uncertain rows
7) Best practices
- Use PDFs directly when available; don’t convert to images unless necessary.
- Keep page coordinates so you can reconstruct tables accurately.
- Use bank-specific templates if the format is consistent.
- Add validation rules for balances and transaction counts.
- Store the raw OCR output for debugging and audits.
- Review low-confidence extractions manually.
8) If you want higher accuracy
Consider using:
- a document AI service trained on forms/tables
- a custom model for your bank statement formats
- post-processing with an LLM to help normalize messy OCR output, but only after OCR extraction
9) Minimal architecture
- Upload PDF/image
- OCR API extracts text + tables
- Parser maps fields into schema
- Validator checks consistency
- Store structured transactions in database
If you want, I can give you:
- a Python example using AWS Textract / Google Document AI / Azure, or
- a regex-based parser for bank statements, or
- a full schema for bank statement extraction.