Prompt
How do I get structured data out of HTML tables and product pages?
Latest observation
To get structured data out of HTML tables and product pages, the usual workflow is:
- Fetch the page HTML
- Parse the DOM
- Identify the relevant elements
- Extract fields into a structured format like JSON, CSV, or a DataFrame
- Clean/normalize the values
1) Extracting data from HTML tables
If the data is in a normal <table>, this is often the easiest case.
Python example with pandas
import pandas as pd
tables = pd.read_html("https://example.com/page-with-table")
df = tables[0] # first table
print(df.head())
This works well when:
- the table is actually rendered in HTML
- the table has simple rows/columns
- you want a quick dataframe
Python example with BeautifulSoup
import requests
from bs4 import BeautifulSoup
url = "https://example.com/page-with-table"
html = requests.get(url).text
soup = BeautifulSoup(html, "html.parser")
table = soup.find("table")
rows = []
headers = [th.get_text(strip=True) for th in table.find_all("th")]
for tr in table.find_all("tr")[1:]:
cells = [td.get_text(strip=True) for td in tr.find_all(["td", "th"])]
if cells:
rows.append(dict(zip(headers, cells)))
print(rows)
Use this if:
- you need more control
- the table is messy
- you need to handle rowspans/colspans manually
2) Extracting data from product pages
Product pages often have data spread across many HTML elements rather than in a table.
Common fields:
- product name
- price
- currency
- SKU / model number
- description
- ratings
- availability
- image URLs
- category
Approach
Inspect the page HTML and locate:
- headings (
<h1>,<h2>) - price elements (
<span class="price">) - metadata tags
- embedded JSON-LD
- script data
- structured attributes like
data-*
Python example with BeautifulSoup
import requests
from bs4 import BeautifulSoup
url = "https://example.com/product-page"
html = requests.get(url).text
soup = BeautifulSoup(html, "html.parser")
name = soup.find("h1").get_text(strip=True)
price = soup.select_one(".price").get_text(strip=True)
description = soup.select_one(".description").get_text(" ", strip=True)
data = {
"name": name,
"price": price,
"description": description
}
print(data)
3) Prefer JSON-LD / embedded structured data when available
Many product pages include schema.org JSON-LD, which is often the cleanest source.
Look for:
<script type="application/ld+json">
{
"@context": "https://schema.org/",
"@type": "Product",
"name": "Example Product",
"offers": {
"@type": "Offer",
"price": "19.99",
"priceCurrency": "USD"
}
}
</script>
Python example
import requests
from bs4 import BeautifulSoup
import json
url = "https://example.com/product-page"
html = requests.get(url).text
soup = BeautifulSoup(html, "html.parser")
for script in soup.find_all("script", type="application/ld+json"):
try:
data = json.loads(script.string)
print(data)
except Exception:
pass
This is great because it often already contains:
- product name
- price
- currency
- availability
- ratings
- brand
- images
4) When pages are rendered by JavaScript
If the HTML source doesn’t contain the content, the page may be rendered dynamically.
In that case:
- use Playwright or Selenium
- or find the underlying API endpoint used by the page
Playwright example
from playwright.sync_api import sync_playwright
with sync_playwright() as p:
browser = p.chromium.launch()
page = browser.new_page()
page.goto("https://example.com/product-page")
page.wait_for_load_state("networkidle")
html = page.content()
browser.close()
Then parse html with BeautifulSoup.
5) Cleaning and normalization
Raw extracted data often needs cleanup:
- remove currency symbols
- convert strings to numbers
- handle missing values
- normalize dates
- standardize units
Example
price_text = "$19.99"
price = float(price_text.replace("$", "").replace(",", ""))
6) If the site is complex, inspect network/API calls
Sometimes the best source is not HTML at all, but a hidden API returning JSON.
Use browser dev tools:
- open Network
- reload page
- look for
fetch,xhr, or API responses - extract data directly from the API if permitted
This is often more reliable than scraping rendered HTML.
7) Important considerations
- Check the site’s terms of service
- Respect robots.txt
- Rate-limit requests
- Use headers / sessions where appropriate
- Be prepared for changes in page structure
Best choice by scenario
- Simple table →
pandas.read_html - Static product page →
requests + BeautifulSoup - Product page with embedded JSON-LD → parse JSON-LD first
- JS-rendered page → Playwright/Selenium
- Highly dynamic site → inspect API calls
If you want, I can show you:
- a generic scraper template for tables and product pages, or
- how to scrape a specific URL if you share one.