Prompt
What should I use to turn scraped pages into structured JSON?
Latest observation
Oct 4, 2026
OpenAI APIWeb search: off
For turning scraped web pages into structured JSON, the best tool depends on how consistent the pages are and whether you want rule-based extraction or AI-based extraction.
Good options
1. BeautifulSoup / lxml + Pydantic
Best if the page structure is fairly consistent.
- Use when: you know the HTML layout
- Pros: fast, cheap, predictable
- Workflow: scrape HTML → parse fields with CSS/XPath → validate into JSON with Pydantic
Example stack:
requestsorhttpxBeautifulSouporlxmlPydanticfor schema validationjsonfor output
2. Scrapy
Best if you're scraping many pages at scale.
- Use when: you need crawling, pagination, retries, pipelines
- Pros: robust, production-friendly
- Workflow: scrape → parse in spider → export JSON
3. Playwright / Selenium + parser
Best for JavaScript-heavy sites.
- Use when: content is rendered dynamically
- Pros: can render page before extraction
- Workflow: render with browser automation → parse HTML → JSON
4. LLM-based extraction
Best if page layouts vary a lot or you want flexible extraction from messy pages.
- Use when: HTML is inconsistent or unstructured
- Pros: less brittle than hand-written selectors
- Cons: slower, more expensive, may need validation
Tools/examples:
- OpenAI / other LLMs with a JSON schema
instructor,outlines, or structured output APIsUnstructuredfor document-like extraction
My recommendation
- If you control the scrape and pages are predictable: use BeautifulSoup/lxml + Pydantic
- If it's a large crawl: use Scrapy
- If pages are JS-rendered: use Playwright + lxml/BeautifulSoup
- If page formats vary a lot: use LLM extraction with schema validation
A practical pipeline
- Fetch page
- Clean HTML
- Extract candidate fields
- Normalize types
- Validate against schema
- Serialize to JSON
If you want, I can suggest a specific Python stack for your use case or show a sample extractor that outputs JSON.
Brands