Prompt
How do I turn unstructured web pages into a clean dataset?
Latest observation
Turning unstructured web pages into a clean dataset usually means building a pipeline like:
- Find the right pages
- Extract the useful content
- Normalize it into fields
- Clean and validate the data
- Store it in a structured format
Here’s a practical approach.
1) Define the target schema first
Before scraping anything, decide what columns you want.
Example for product pages:
namepricebrandratingreview_counturlscraped_at
If you don’t define a schema early, you’ll end up with messy, inconsistent output.
2) Collect the pages
Ways to gather pages:
- Static sites:
requests+BeautifulSoup - JavaScript-heavy sites:
Playwright,Selenium, orScrapy + browser rendering - Large-scale crawling:
Scrapy,Heritrix, custom crawler
Good practice:
- Respect
robots.txt - Add rate limiting
- Identify yourself with a user agent
- Cache pages so you don’t re-download unnecessarily
3) Extract the relevant content
Web pages are messy because of navigation, ads, popups, and repeated boilerplate.
Common extraction methods
A. CSS/XPath selectors
Best when the HTML structure is consistent.
from bs4 import BeautifulSoup
import requests
html = requests.get("https://example.com").text
soup = BeautifulSoup(html, "html.parser")
title = soup.select_one("h1").get_text(strip=True)
price = soup.select_one(".price").get_text(strip=True)
B. Structured data
Many sites embed JSON-LD or microdata.
Look for:
<script type="application/ld+json">schema.orgmarkup
This is often cleaner than parsing visible HTML.
C. Content extraction libraries
If you want article text:
trafilaturareadability-lxmlnewspaper3k
These remove boilerplate and keep main content.
D. LLM-assisted extraction
For highly variable pages, you can use an LLM to convert page content into a fixed schema.
Example prompt:
Extract
title,date,author, andmain_textfrom this HTML and return valid JSON matching this schema.
Useful when:
- Page layouts vary
- Fields are not in predictable HTML
- You need flexible extraction
But still validate the output carefully.
4) Normalize the data
Raw scraped data often contains inconsistent formats.
Examples:
- Prices like
$12.99,12,99 EUR,USD 12.99 - Dates like
Jan 3, 2025or2025-01-03 - Names with extra whitespace
- Numbers with commas
Normalize them into standard types:
- strings -> trimmed
- prices -> numeric + currency
- dates -> ISO 8601
- booleans -> true/false
- lists -> proper arrays or separate table rows
Example:
import re
from datetime import datetime
def parse_price(text):
match = re.search(r"[\d,.]+", text)
return float(match.group().replace(",", "")) if match else None
def parse_date(text):
return datetime.strptime(text, "%b %d, %Y").date().isoformat()
5) Clean the dataset
Typical cleaning steps:
- Remove duplicates
- Handle missing values
- Standardize categories
- Filter out bad rows
- Deduplicate by URL or ID
- Strip HTML tags and whitespace
- Fix encoding issues
Example checks:
- Is
pricenumeric? - Is
urlvalid? - Is
nameempty? - Are there duplicate records?
6) Validate the output
Don’t assume extraction worked.
Useful validation:
- Count rows vs expected pages
- Check null rates per column
- Verify sample records manually
- Check outliers
- Assert schema types
Tools:
pandaspanderagreat_expectations
Example with pandas:
import pandas as pd
df = pd.read_json("data.jsonl", lines=True)
print(df.isna().mean())
print(df.head())
7) Store it in a usable format
Good options:
- CSV: simple, human-readable
- JSONL: great for nested or row-based records
- Parquet: best for analytics and large datasets
- SQL database: good for querying and updates
Recommended:
- Use JSONL during scraping
- Convert to Parquet or a database after cleaning
8) Handle difficult cases
Dynamic pages
Use a browser automation tool:
- Playwright is often the best choice today.
Sites with anti-bot measures
- Slow down requests
- Rotate user agents responsibly
- Use browser rendering if needed
- Avoid anything that violates site rules
Mixed layouts
Use:
- Page-type classifiers
- Multiple extraction rules
- LLM fallback for edge cases
Noisy text extraction
Try boilerplate removal libraries or main-content extractors.
9) A simple practical pipeline
A common workflow:
- Crawl URLs
- Download HTML
- Extract fields with selectors or JSON-LD
- Normalize and clean fields
- Validate schema
- Export to CSV/Parquet
10) Example end-to-end structure
For a scraping project, you might organize it like this:
project/
crawl.py
extract.py
clean.py
validate.py
data/
raw/
processed/
crawl.pysaves raw HTMLextract.pyparses HTML into recordsclean.pystandardizes valuesvalidate.pychecks quality
When to use an LLM
Use an LLM when:
- Page structures vary a lot
- You need to extract semantic fields from messy content
- You have moderate data volume
Avoid relying on it alone when:
- You need strict reproducibility
- You’re processing millions of pages
- You need guaranteed exact parsing
Best pattern:
- Deterministic parsing first
- LLM only as fallback
If you want, I can also show you:
- a Python scraping template,
- a Playwright example for JS-heavy pages, or
- a pipeline that uses an LLM to extract structured JSON from HTML.