\n```\n\n### Python example\n```python\nimport requests\nfrom bs4 import BeautifulSoup\nimport json\n\nurl = \"https://example.com/product-page\"\nhtml = requests.get(url).text\nsoup = BeautifulSoup(html, \"html.parser\")\n\nfor script in soup.find_all(\"script\", type=\"application/ld+json\"):\n try:\n data = json.loads(script.string)\n print(data)\n except Exception:\n pass\n```\n\nThis is great because it often already contains:\n- product name\n- price\n- currency\n- availability\n- ratings\n- brand\n- images\n\n---\n\n## 4) When pages are rendered by JavaScript\n\nIf the HTML source doesn’t contain the content, the page may be rendered dynamically.\n\nIn that case:\n- use **Playwright** or **Selenium**\n- or find the underlying API endpoint used by the page\n\n### Playwright example\n```python\nfrom playwright.sync_api import sync_playwright\n\nwith sync_playwright() as p:\n browser = p.chromium.launch()\n page = browser.new_page()\n page.goto(\"https://example.com/product-page\")\n page.wait_for_load_state(\"networkidle\")\n html = page.content()\n browser.close()\n```\n\nThen parse `html` with BeautifulSoup.\n\n---\n\n## 5) Cleaning and normalization\n\nRaw extracted data often needs cleanup:\n- remove currency symbols\n- convert strings to numbers\n- handle missing values\n- normalize dates\n- standardize units\n\n### Example\n```python\nprice_text = \"$19.99\"\nprice = float(price_text.replace(\"$\", \"\").replace(\",\", \"\"))\n```\n\n---\n\n## 6) If the site is complex, inspect network/API calls\n\nSometimes the best source is not HTML at all, but a hidden API returning JSON.\n\nUse browser dev tools:\n- open **Network**\n- reload page\n- look for `fetch`, `xhr`, or API responses\n- extract data directly from the API if permitted\n\nThis is often more reliable than scraping rendered HTML.\n\n---\n\n## 7) Important considerations\n\n- Check the site’s **terms of service**\n- Respect **robots.txt**\n- Rate-limit requests\n- Use headers / sessions where appropriate\n- Be prepared for changes in page structure\n\n---\n\n## Best choice by scenario\n\n- **Simple table** → `pandas.read_html`\n- **Static product page** → `requests + BeautifulSoup`\n- **Product page with embedded JSON-LD** → parse JSON-LD first\n- **JS-rendered page** → Playwright/Selenium\n- **Highly dynamic site** → inspect API calls\n\nIf you want, I can show you:\n1. a **generic scraper template** for "}}}

Prompt

How do I get structured data out of HTML tables and product pages?

Technology / API Platforms1 observationLast seen Oct 4, 2026

Latest observation

Oct 4, 2026OpenAI APIWeb search: off

To get structured data out of HTML tables and product pages, the usual workflow is:

  1. Fetch the page HTML
  2. Parse the DOM
  3. Identify the relevant elements
  4. Extract fields into a structured format like JSON, CSV, or a DataFrame
  5. Clean/normalize the values

1) Extracting data from HTML tables

If the data is in a normal <table>, this is often the easiest case.

Python example with pandas

import pandas as pd

tables = pd.read_html("https://example.com/page-with-table")
df = tables[0]  # first table
print(df.head())

This works well when:

  • the table is actually rendered in HTML
  • the table has simple rows/columns
  • you want a quick dataframe

Python example with BeautifulSoup

import requests
from bs4 import BeautifulSoup

url = "https://example.com/page-with-table"
html = requests.get(url).text
soup = BeautifulSoup(html, "html.parser")

table = soup.find("table")
rows = []

headers = [th.get_text(strip=True) for th in table.find_all("th")]
for tr in table.find_all("tr")[1:]:
    cells = [td.get_text(strip=True) for td in tr.find_all(["td", "th"])]
    if cells:
        rows.append(dict(zip(headers, cells)))

print(rows)

Use this if:

  • you need more control
  • the table is messy
  • you need to handle rowspans/colspans manually

2) Extracting data from product pages

Product pages often have data spread across many HTML elements rather than in a table.

Common fields:

  • product name
  • price
  • currency
  • SKU / model number
  • description
  • ratings
  • availability
  • image URLs
  • category

Approach

Inspect the page HTML and locate:

  • headings (<h1>, <h2>)
  • price elements (<span class="price">)
  • metadata tags
  • embedded JSON-LD
  • script data
  • structured attributes like data-*

Python example with BeautifulSoup

import requests
from bs4 import BeautifulSoup

url = "https://example.com/product-page"
html = requests.get(url).text
soup = BeautifulSoup(html, "html.parser")

name = soup.find("h1").get_text(strip=True)
price = soup.select_one(".price").get_text(strip=True)
description = soup.select_one(".description").get_text(" ", strip=True)

data = {
    "name": name,
    "price": price,
    "description": description
}

print(data)

3) Prefer JSON-LD / embedded structured data when available

Many product pages include schema.org JSON-LD, which is often the cleanest source.

Look for:

<script type="application/ld+json">
{
  "@context": "https://schema.org/",
  "@type": "Product",
  "name": "Example Product",
  "offers": {
    "@type": "Offer",
    "price": "19.99",
    "priceCurrency": "USD"
  }
}
</script>

Python example

import requests
from bs4 import BeautifulSoup
import json

url = "https://example.com/product-page"
html = requests.get(url).text
soup = BeautifulSoup(html, "html.parser")

for script in soup.find_all("script", type="application/ld+json"):
    try:
        data = json.loads(script.string)
        print(data)
    except Exception:
        pass

This is great because it often already contains:

  • product name
  • price
  • currency
  • availability
  • ratings
  • brand
  • images

4) When pages are rendered by JavaScript

If the HTML source doesn’t contain the content, the page may be rendered dynamically.

In that case:

  • use Playwright or Selenium
  • or find the underlying API endpoint used by the page

Playwright example

from playwright.sync_api import sync_playwright

with sync_playwright() as p:
    browser = p.chromium.launch()
    page = browser.new_page()
    page.goto("https://example.com/product-page")
    page.wait_for_load_state("networkidle")
    html = page.content()
    browser.close()

Then parse html with BeautifulSoup.


5) Cleaning and normalization

Raw extracted data often needs cleanup:

  • remove currency symbols
  • convert strings to numbers
  • handle missing values
  • normalize dates
  • standardize units

Example

price_text = "$19.99"
price = float(price_text.replace("$", "").replace(",", ""))

6) If the site is complex, inspect network/API calls

Sometimes the best source is not HTML at all, but a hidden API returning JSON.

Use browser dev tools:

  • open Network
  • reload page
  • look for fetch, xhr, or API responses
  • extract data directly from the API if permitted

This is often more reliable than scraping rendered HTML.


7) Important considerations

  • Check the site’s terms of service
  • Respect robots.txt
  • Rate-limit requests
  • Use headers / sessions where appropriate
  • Be prepared for changes in page structure

Best choice by scenario

  • Simple table → pandas.read_html
  • Static product page → requests + BeautifulSoup
  • Product page with embedded JSON-LD → parse JSON-LD first
  • JS-rendered page → Playwright/Selenium
  • Highly dynamic site → inspect API calls

If you want, I can show you:

  1. a generic scraper template for tables and product pages, or
  2. how to scrape a specific URL if you share one.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.