Prompt

How do I turn unstructured web pages into a clean CSV file?

Technology / API Platforms1 observationLast seen Oct 4, 2026

Latest observation

Oct 4, 2026OpenAI APIWeb search: off

To turn unstructured web pages into a clean CSV, the basic pipeline is:

  1. Fetch the page
  2. Parse the HTML
  3. Extract the fields you want
  4. Clean/normalize the data
  5. Write rows to CSV

1) Inspect the page first

Open the page in your browser and use View Source or Developer Tools to find:

  • repeating blocks (cards, list items, table rows)
  • CSS classes / tags around the data
  • whether the content is loaded by JavaScript

If the data is in a normal HTML table, extraction is easy. If it’s rendered by JavaScript, you may need a browser automation tool.


2) Choose the right tool

If the page is static HTML

Use:

  • Python + requests + BeautifulSoup
  • pandas.read_html for tables
  • lxml for faster parsing

If the page is dynamic / JavaScript-rendered

Use:

  • Playwright
  • Selenium
  • Or find the site’s underlying API calls in DevTools → Network

3) Example: static page with repeating items

import requests
from bs4 import BeautifulSoup
import csv

url = "https://example.com/products"
html = requests.get(url, timeout=30).text
soup = BeautifulSoup(html, "html.parser")

rows = []
for item in soup.select(".product-card"):
    name = item.select_one(".name").get_text(strip=True)
    price = item.select_one(".price").get_text(strip=True)
    link = item.select_one("a")["href"]

    rows.append({
        "name": name,
        "price": price,
        "link": link
    })

with open("output.csv", "w", newline="", encoding="utf-8") as f:
    writer = csv.DictWriter(f, fieldnames=["name", "price", "link"])
    writer.writeheader()
    writer.writerows(rows)

4) Example: table on the page

import pandas as pd

tables = pd.read_html("https://example.com/stats")
tables[0].to_csv("output.csv", index=False)

If there are multiple tables, inspect tables[0], tables[1], etc.


5) Cleaning the data

Web data is often messy. Common cleanup steps:

  • remove extra whitespace
  • convert currency strings to numbers
  • standardize dates
  • handle missing values
  • resolve relative URLs

Example:

import re

def clean_price(text):
    text = text.replace("$", "").replace(",", "").strip()
    return float(text) if text else None

def clean_text(text):
    return re.sub(r"\s+", " ", text).strip()

6) If the page is JavaScript-rendered

Use Playwright to load the page fully:

from playwright.sync_api import sync_playwright
from bs4 import BeautifulSoup
import csv

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    page = browser.new_page()
    page.goto("https://example.com/products", wait_until="networkidle")
    html = page.content()
    browser.close()

soup = BeautifulSoup(html, "html.parser")
# parse soup as usual

7) Handle pagination

If the site has many pages, loop through them:

  • page numbers in URL
  • “next” buttons
  • API endpoints

Example:

for page_num in range(1, 6):
    url = f"https://example.com/products?page={page_num}"
    # fetch and parse each page

8) Best practices

  • Check the site’s robots.txt and terms of service
  • Add delays between requests if needed
  • Set a User-Agent header
  • Expect broken/missing data and code defensively

Example headers:

headers = {
    "User-Agent": "Mozilla/5.0"
}
requests.get(url, headers=headers)

9) When the page is “really” unstructured

If the content is free-form text, you may need:

  • regex
  • natural language processing
  • manual rules
  • an LLM-assisted extraction step

For example, you can extract fields from article text like:

  • title
  • date
  • location
  • price
  • contact info

Then assemble those into rows for CSV.


If you want, I can give you:

  1. a Python template for scraping pages into CSV, or
  2. a method for your specific website if you share a URL or sample HTML.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as Circuit. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.