Prompt
How do I extract fields from documents with an LLM API?
Latest observation
Oct 8, 2026
OpenAI APIWeb search: off
To extract fields from documents with an LLM API, the usual pattern is:
- Convert the document to text
- Define the fields you want
- Prompt the model to return structured output
- Parse and validate the result
- Handle edge cases and retries
1) Extract text from the document
If your input is:
- PDF / Word / image: use OCR or document parsing first
- plain text / HTML / JSON: you may be able to send it directly
Common tools:
- PDFs:
pymupdf,pdfplumber,unstructured - Images: OCR like
tesseract, cloud OCR, or a vision-capable model - Scanned docs: OCR first, then LLM
2) Decide the schema
Be explicit about what fields you want, for example:
{
"invoice_number": "",
"date": "",
"vendor_name": "",
"total_amount": "",
"currency": ""
}
3) Prompt the LLM for structured output
Ask for only JSON and specify the schema.
Example prompt:
Extract the following fields from the document text:
- invoice_number
- date
- vendor_name
- total_amount
- currency
Return valid JSON only. If a field is missing, use null.
4) Example API call
Here’s a generic example in Python:
import requests
import json
document_text = open("doc.txt").read()
prompt = f"""
Extract these fields from the document:
- invoice_number
- date
- vendor_name
- total_amount
- currency
Return valid JSON only. If missing, use null.
Document:
{document_text}
"""
response = requests.post(
"https://api.your-llm.com/v1/chat/completions",
headers={"Authorization": "Bearer YOUR_API_KEY"},
json={
"model": "your-model",
"messages": [
{"role": "system", "content": "You extract structured data from documents."},
{"role": "user", "content": prompt}
],
"temperature": 0
}
)
print(response.json())
5) Prefer schema-constrained output if available
If your API supports structured outputs, function calling, or JSON schema mode, use that. It reduces parsing errors a lot.
Example idea:
- Define a JSON schema for the fields
- The model must fill it
- Your code validates the result
6) Validate the output
Check:
- JSON parses correctly
- Required fields exist
- Dates are in a standard format
- Amounts are numeric
- Confidence or source text is reasonable
If validation fails, retry with a stricter prompt or fallback logic.
7) For better accuracy
A few practical tips:
- Use
temperature=0 - Keep the prompt short and specific
- Provide field definitions and examples
- Include a few sample documents + expected outputs if possible
- Chunk long documents and merge results if needed
- Ask the model to cite the exact text span it used
Example with source evidence:
{
"invoice_number": "INV-12345",
"invoice_number_source": "Invoice #: INV-12345"
}
8) If you have many document types
Use a pipeline:
- classify document type
- choose a schema for that type
- extract fields
- validate against rules
If you want, I can give you:
- a Python example using OpenAI-style structured outputs,
- a JSON schema template for document extraction, or
- a full invoice/receipt extraction pipeline.