Prompt
I'm building an AI data pipeline from public websites and need structured output, not raw HTML
Latest observation
Sep 24, 2026
OpenAI APIWeb search: off
If you’re pulling from public websites and want structured data instead of raw HTML, the best approach is to combine targeted extraction with a fixed output schema.
Recommended pipeline
-
Fetch page content
- Use
requests,httpx, or a crawler. - Respect robots.txt, rate limits, and site terms.
- Use
-
Clean/parse the page
- Extract visible text and relevant DOM sections.
- Libraries:
BeautifulSoup,lxml,readability-lxml,trafilatura.
-
Map page content to a schema
- Define a JSON schema for the data you need.
- Example fields: title, author, published_date, price, product_name, summary, source_url.
-
Validate structured output
- Ensure each record matches your schema.
- Use
pydantic,jsonschema, or similar.
-
Store in a structured format
- JSONL for pipelines, PostgreSQL for queryable storage, Parquet for analytics.
Example structured schema
{
"source_url": "https://example.com/article",
"title": "Example Article",
"author": "Jane Doe",
"published_date": "2026-09-24",
"summary": "Short summary of the page.",
"tags": ["news", "ai"]
}
Practical extraction options
1) DOM selector-based extraction
Best when sites are consistent.
from bs4 import BeautifulSoup
import requests
url = "https://example.com"
html = requests.get(url, timeout=20).text
soup = BeautifulSoup(html, "lxml")
data = {
"title": soup.select_one("h1").get_text(strip=True) if soup.select_one("h1") else None,
"source_url": url
}
print(data)
2) Main-content extraction
Best for articles/blog posts.
import trafilatura
downloaded = trafilatura.fetch_url("https://example.com/article")
text = trafilatura.extract(downloaded, include_comments=False, include_tables=True)
print(text)
3) LLM-to-JSON extraction
Best when page structure varies and you need flexible parsing.
- Feed cleaned text/HTML fragments to the model.
- Ask for strict JSON matching your schema.
- Validate afterward.
Tips for reliable structured output
- Prefer site-specific parsers for high-value sources.
- Keep a per-site extraction config if you’re doing this at scale.
- Normalize fields:
- dates to ISO-8601
- prices to numeric + currency
- lists to arrays
- Deduplicate records using canonical URLs and content hashes.
- Log extraction confidence and missing fields.
Suggested output format for pipelines
For most data pipelines, use JSON Lines:
{"source_url":"...","title":"...","author":"..."}
{"source_url":"...","title":"...","author":"..."}
If you want, I can help you with one of these next:
- a generic Python scraper that outputs JSONL,
- a site-specific extraction template, or
- an LLM-based HTML-to-structured-data pipeline.