Prompt
building a document extraction pipeline with llm api
Latest observation
Here’s a practical blueprint for building a document extraction pipeline with an LLM API.
1) Define the extraction target
Be explicit about what you want from each document.
Examples:
- Invoice fields:
invoice_number,vendor,date,subtotal,tax,total - Contract fields:
parties,effective_date,termination_clause,jurisdiction - Resume fields:
name,email,experience,skills
Create a schema first. Prefer structured output like JSON.
2) Typical pipeline architecture
A. Ingest
- Receive files: PDF, DOCX, images, emails, etc.
- Store originals in object storage
- Track metadata: source, upload time, document type, customer id
B. Preprocess
- Convert to text when possible:
- PDF text extraction
- OCR for scanned PDFs/images
- DOCX parsing
- Normalize:
- remove headers/footers if noisy
- preserve page numbers
- split into chunks if long
C. Classify document type
Use a lightweight classifier or LLM prompt to determine:
- invoice
- contract
- receipt
- tax form
- etc.
This helps route to the right extraction schema.
D. Extract with LLM
Send:
- document text or chunked text
- extraction instructions
- schema
- examples if needed
Ask for:
- structured JSON only
- confidence fields if useful
- citations / source spans / page numbers if possible
E. Validate
After LLM output:
- JSON schema validation
- type checks
- required field checks
- format checks for dates, currency, IDs
- cross-field consistency checks
F. Post-process
- Normalize dates
- Convert currencies
- Clean names/addresses
- Merge chunk-level extractions
- Deduplicate repeated fields
G. Human review for low confidence
Route uncertain or invalid outputs to review UI.
H. Store results
Store:
- extracted JSON
- confidence scores
- source references
- validation errors
- model/version used
3) Recommended extraction pattern
Best practice: use structured outputs
Instead of asking for free-form text, request a strict schema.
Example output:
{
"invoice_number": "INV-1042",
"vendor": "Acme Corp",
"invoice_date": "2025-08-14",
"subtotal": 1200.00,
"tax": 96.00,
"total": 1296.00,
"currency": "USD"
}
If your LLM API supports it, use:
- JSON schema response format
- function calling / tool calling
- constrained decoding if available
4) Chunking strategy for long documents
For long docs:
- split by page or section
- extract from each chunk
- merge results
Important:
- preserve page/section metadata
- avoid splitting tables mid-row
- overlap chunks slightly if needed
For multi-page invoices/contracts, often:
- first pass: detect relevant pages
- second pass: extract fields from those pages only
5) Prompt design tips
Good prompt ingredients
- Clear task
- Exact schema
- Field definitions
- Rules for missing data
- Examples of valid/invalid output
Example prompt
Extract the following fields from the document:
- invoice_number: string
- vendor_name: string
- invoice_date: YYYY-MM-DD
- total_amount: number
- currency: ISO 4217 code
Rules:
- Return only valid JSON.
- If a field is missing, use null.
- Do not guess.
- Use the document text only.
6) Handling OCR noise and scanned docs
If documents are scanned:
- run OCR first
- keep OCR confidence per word/page if possible
- pass raw OCR text to the LLM
- consider a layout-aware OCR tool
For tables:
- preserve row/column structure where possible
- extract tables separately if your pipeline supports it
7) Quality controls
Add multiple layers:
Deterministic checks
- regex for dates, IDs, emails
- totals match subtotal + tax
- required fields present
Model-based checks
- ask LLM to verify its own output against the source
- use a second pass for validation
Human-in-the-loop
- route edge cases:
- low OCR confidence
- missing required fields
- inconsistent totals
- ambiguous documents
8) Metrics to track
Measure pipeline quality with:
- field-level precision/recall/F1
- exact match rate
- document-level success rate
- invalid JSON rate
- human review rate
- cost per document
- latency per document
9) Common failure modes
- hallucinated fields
- wrong dates/currencies
- duplicate values from repeated headers
- OCR errors
- tables losing alignment
- long-document context overflow
- model variability across versions
Mitigations:
- strict schema
- null instead of guessing
- validation rules
- chunking
- citations/page references
- fallback to human review
10) Minimal implementation flow
- Upload document
- OCR/text extraction
- Detect document type
- Call LLM with schema
- Validate JSON
- Normalize and store
- Send exceptions to review
11) Simple pseudo-code
doc = load_document(file)
text = ocr_or_extract_text(doc)
doc_type = classify_doc_type(text)
schema = schemas[doc_type]
prompt = build_prompt(text, schema)
result = llm.extract(prompt, response_format="json")
validated = validate_against_schema(result, schema)
if not validated.ok:
send_to_review(doc, result, validated.errors)
else:
save_extracted_data(doc.id, result)
12) Suggested stack
Depending on your needs:
Parsing/OCR
- pdfplumber, pymupdf
- Tesseract
- AWS Textract
- Google Document AI
- Azure Form Recognizer
LLM orchestration
- OpenAI / Anthropic / Gemini APIs
- LangChain / LlamaIndex if you need orchestration
- Pydantic / JSON Schema for validation
Storage
- S3/GCS/Azure Blob for files
- Postgres for extracted records
- Elasticsearch/OpenSearch if you need search
13) Strong recommendation
If the documents are semi-structured (invoices, receipts, forms), combine:
- OCR/layout extraction tools
- LLM only for reasoning and normalization
This is usually more reliable and cheaper than using the LLM alone.
If you want, I can give you one of these next:
- a reference architecture diagram
- a Python implementation example
- a prompt template + JSON schema
- a production-ready pipeline design for invoices/contracts