Prompt

How do I extract fields from documents with an LLM API?

Artificial Intelligence / AI Platforms1 observationLast seen Oct 8, 2026

Latest observation

Oct 8, 2026OpenAI APIWeb search: off

To extract fields from documents with an LLM API, the usual pattern is:

  1. Convert the document to text
  2. Define the fields you want
  3. Prompt the model to return structured output
  4. Parse and validate the result
  5. Handle edge cases and retries

1) Extract text from the document

If your input is:

  • PDF / Word / image: use OCR or document parsing first
  • plain text / HTML / JSON: you may be able to send it directly

Common tools:

  • PDFs: pymupdf, pdfplumber, unstructured
  • Images: OCR like tesseract, cloud OCR, or a vision-capable model
  • Scanned docs: OCR first, then LLM

2) Decide the schema

Be explicit about what fields you want, for example:

{
  "invoice_number": "",
  "date": "",
  "vendor_name": "",
  "total_amount": "",
  "currency": ""
}

3) Prompt the LLM for structured output

Ask for only JSON and specify the schema.

Example prompt:

Extract the following fields from the document text:

  • invoice_number
  • date
  • vendor_name
  • total_amount
  • currency
    Return valid JSON only. If a field is missing, use null.

4) Example API call

Here’s a generic example in Python:

import requests
import json

document_text = open("doc.txt").read()

prompt = f"""
Extract these fields from the document:
- invoice_number
- date
- vendor_name
- total_amount
- currency

Return valid JSON only. If missing, use null.

Document:
{document_text}
"""

response = requests.post(
    "https://api.your-llm.com/v1/chat/completions",
    headers={"Authorization": "Bearer YOUR_API_KEY"},
    json={
        "model": "your-model",
        "messages": [
            {"role": "system", "content": "You extract structured data from documents."},
            {"role": "user", "content": prompt}
        ],
        "temperature": 0
    }
)

print(response.json())

5) Prefer schema-constrained output if available

If your API supports structured outputs, function calling, or JSON schema mode, use that. It reduces parsing errors a lot.

Example idea:

  • Define a JSON schema for the fields
  • The model must fill it
  • Your code validates the result

6) Validate the output

Check:

  • JSON parses correctly
  • Required fields exist
  • Dates are in a standard format
  • Amounts are numeric
  • Confidence or source text is reasonable

If validation fails, retry with a stricter prompt or fallback logic.

7) For better accuracy

A few practical tips:

  • Use temperature=0
  • Keep the prompt short and specific
  • Provide field definitions and examples
  • Include a few sample documents + expected outputs if possible
  • Chunk long documents and merge results if needed
  • Ask the model to cite the exact text span it used

Example with source evidence:

{
  "invoice_number": "INV-12345",
  "invoice_number_source": "Invoice #: INV-12345"
}

8) If you have many document types

Use a pipeline:

  • classify document type
  • choose a schema for that type
  • extract fields
  • validate against rules

If you want, I can give you:

  1. a Python example using OpenAI-style structured outputs,
  2. a JSON schema template for document extraction, or
  3. a full invoice/receipt extraction pipeline.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.