Prompt
I'm building an internal research tool and need structured extraction from web pages
Latest observation
Oct 4, 2026
OpenAI APIWeb search: off
Sure — I can help with structured extraction from web pages.
A good pattern is:
- Fetch the page content
- Extract the relevant fields
- Normalize into a schema
- Return JSON
If you want, I can help with any of these in a few ways:
Options I can provide
- A scraping/extraction plan for a specific site
- Python code using
requests,BeautifulSoup,lxml, orplaywright - JavaScript/Node code using
fetch,cheerio, orplaywright - A JSON schema for the fields you need
- An LLM-based extraction prompt that converts page text/HTML into structured output
- A hybrid approach: deterministic parsing first, LLM cleanup second
Typical extraction workflow
For each page, you might collect:
urltitlepublished_dateauthorsummarymain_contententities/tagstables/list itemssource_metadata(site name, crawl time, language)
Example output format
{
"url": "https://example.com/article",
"title": "Example Article",
"published_date": "2026-10-04",
"author": "Jane Doe",
"summary": "Short summary here.",
"main_content": "Extracted body text...",
"tags": ["research", "news"],
"source_metadata": {
"site": "example.com",
"extracted_at": "2026-10-04T12:00:00Z"
}
}
If you want an LLM extraction prompt
A common prompt pattern is:
- Provide the page HTML or cleaned text
- Specify the schema explicitly
- Require valid JSON only
- Instruct the model to use
nullwhen a field is missing
Example:
Extract the following fields from the page content and return valid JSON only:
- title
- author
- published_date
- summary
- key_points
Rules:
- If a field is missing, use null.
- Do not invent data.
- Output only JSON.
If you share:
- the website/page type,
- the fields you need, and
- your preferred stack (Python/JS/etc.),
I can draft a concrete extractor for you.