Prompt
structured data extraction from web pages
Latest observation
Oct 4, 2026
OpenAI APIWeb search: off
Structured data extraction from web pages means turning information on a webpage into organized fields you can store, search, analyze, or feed into another system.
Common approaches
-
Direct HTML parsing
- Extract content from the DOM using CSS selectors or XPath.
- Best when page structure is stable.
-
Embedded structured data
- Many pages include:
- JSON-LD
- Microdata
- RDFa
- Often the cleanest source for things like products, articles, events, recipes.
- Many pages include:
-
Browser automation / rendering
- Use Playwright, Puppeteer, Selenium, etc. to handle JavaScript-rendered pages.
- Needed when content is loaded dynamically.
-
Text + LLM-based extraction
- Convert page content to text and ask a model to extract fields into a schema.
- Useful for messy or inconsistent pages, but less deterministic.
Typical pipeline
- Fetch page
- Render if needed
- Detect and parse structured data
- Fall back to DOM extraction
- Normalize fields
- Validate against a schema
- Store in JSON / database / CSV
Example schema
{
"title": "",
"author": "",
"published_date": "",
"price": "",
"currency": "",
"availability": "",
"url": ""
}
Tools often used
- Python: BeautifulSoup, lxml, requests, Scrapy, Playwright
- JavaScript/TypeScript: Cheerio, Playwright, Puppeteer
- Structured data parsers: extruct, schema.org parsers
- Validation: Pydantic, JSON Schema
Example: extracting JSON-LD in Python
import requests
from bs4 import BeautifulSoup
import json
html = requests.get("https://example.com").text
soup = BeautifulSoup(html, "html.parser")
for script in soup.find_all("script", type="application/ld+json"):
try:
data = json.loads(script.string)
print(data)
except Exception:
pass
Best practices
- Prefer structured data already present on the page
- Validate extracted values
- Handle pagination and duplicate records
- Respect robots.txt and site terms
- Add retries, timeouts, and rate limiting
- Log extraction failures for debugging
If you want, I can also provide:
- a complete Python scraper template
- a Playwright-based extractor
- or help design an extraction schema for your use case.