Prompt
How do I turn unstructured web pages into a clean CSV file?
Latest observation
Oct 4, 2026
OpenAI APIWeb search: off
To turn unstructured web pages into a clean CSV, the basic pipeline is:
- Fetch the page
- Parse the HTML
- Extract the fields you want
- Clean/normalize the data
- Write rows to CSV
1) Inspect the page first
Open the page in your browser and use View Source or Developer Tools to find:
- repeating blocks (cards, list items, table rows)
- CSS classes / tags around the data
- whether the content is loaded by JavaScript
If the data is in a normal HTML table, extraction is easy. If it’s rendered by JavaScript, you may need a browser automation tool.
2) Choose the right tool
If the page is static HTML
Use:
- Python + requests + BeautifulSoup
- pandas.read_html for tables
- lxml for faster parsing
If the page is dynamic / JavaScript-rendered
Use:
- Playwright
- Selenium
- Or find the site’s underlying API calls in DevTools → Network
3) Example: static page with repeating items
import requests
from bs4 import BeautifulSoup
import csv
url = "https://example.com/products"
html = requests.get(url, timeout=30).text
soup = BeautifulSoup(html, "html.parser")
rows = []
for item in soup.select(".product-card"):
name = item.select_one(".name").get_text(strip=True)
price = item.select_one(".price").get_text(strip=True)
link = item.select_one("a")["href"]
rows.append({
"name": name,
"price": price,
"link": link
})
with open("output.csv", "w", newline="", encoding="utf-8") as f:
writer = csv.DictWriter(f, fieldnames=["name", "price", "link"])
writer.writeheader()
writer.writerows(rows)
4) Example: table on the page
import pandas as pd
tables = pd.read_html("https://example.com/stats")
tables[0].to_csv("output.csv", index=False)
If there are multiple tables, inspect tables[0], tables[1], etc.
5) Cleaning the data
Web data is often messy. Common cleanup steps:
- remove extra whitespace
- convert currency strings to numbers
- standardize dates
- handle missing values
- resolve relative URLs
Example:
import re
def clean_price(text):
text = text.replace("$", "").replace(",", "").strip()
return float(text) if text else None
def clean_text(text):
return re.sub(r"\s+", " ", text).strip()
6) If the page is JavaScript-rendered
Use Playwright to load the page fully:
from playwright.sync_api import sync_playwright
from bs4 import BeautifulSoup
import csv
with sync_playwright() as p:
browser = p.chromium.launch(headless=True)
page = browser.new_page()
page.goto("https://example.com/products", wait_until="networkidle")
html = page.content()
browser.close()
soup = BeautifulSoup(html, "html.parser")
# parse soup as usual
7) Handle pagination
If the site has many pages, loop through them:
- page numbers in URL
- “next” buttons
- API endpoints
Example:
for page_num in range(1, 6):
url = f"https://example.com/products?page={page_num}"
# fetch and parse each page
8) Best practices
- Check the site’s robots.txt and terms of service
- Add delays between requests if needed
- Set a User-Agent header
- Expect broken/missing data and code defensively
Example headers:
headers = {
"User-Agent": "Mozilla/5.0"
}
requests.get(url, headers=headers)
9) When the page is “really” unstructured
If the content is free-form text, you may need:
- regex
- natural language processing
- manual rules
- an LLM-assisted extraction step
For example, you can extract fields from article text like:
- title
- date
- location
- price
- contact info
Then assemble those into rows for CSV.
If you want, I can give you:
- a Python template for scraping pages into CSV, or
- a method for your specific website if you share a URL or sample HTML.