Prompt

How do I turn unstructured web pages into a clean dataset?

Technology / API Platforms1 observationLast seen Oct 4, 2026

Latest observation

Oct 4, 2026OpenAI APIWeb search: off

Turning unstructured web pages into a clean dataset usually means building a pipeline like:

  1. Find the right pages
  2. Extract the useful content
  3. Normalize it into fields
  4. Clean and validate the data
  5. Store it in a structured format

Here’s a practical approach.


1) Define the target schema first

Before scraping anything, decide what columns you want.

Example for product pages:

  • name
  • price
  • brand
  • rating
  • review_count
  • url
  • scraped_at

If you don’t define a schema early, you’ll end up with messy, inconsistent output.


2) Collect the pages

Ways to gather pages:

  • Static sites: requests + BeautifulSoup
  • JavaScript-heavy sites: Playwright, Selenium, or Scrapy + browser rendering
  • Large-scale crawling: Scrapy, Heritrix, custom crawler

Good practice:

  • Respect robots.txt
  • Add rate limiting
  • Identify yourself with a user agent
  • Cache pages so you don’t re-download unnecessarily

3) Extract the relevant content

Web pages are messy because of navigation, ads, popups, and repeated boilerplate.

Common extraction methods

A. CSS/XPath selectors

Best when the HTML structure is consistent.

from bs4 import BeautifulSoup
import requests

html = requests.get("https://example.com").text
soup = BeautifulSoup(html, "html.parser")

title = soup.select_one("h1").get_text(strip=True)
price = soup.select_one(".price").get_text(strip=True)

B. Structured data

Many sites embed JSON-LD or microdata.

Look for:

  • <script type="application/ld+json">
  • schema.org markup

This is often cleaner than parsing visible HTML.

C. Content extraction libraries

If you want article text:

  • trafilatura
  • readability-lxml
  • newspaper3k

These remove boilerplate and keep main content.

D. LLM-assisted extraction

For highly variable pages, you can use an LLM to convert page content into a fixed schema.

Example prompt:

Extract title, date, author, and main_text from this HTML and return valid JSON matching this schema.

Useful when:

  • Page layouts vary
  • Fields are not in predictable HTML
  • You need flexible extraction

But still validate the output carefully.


4) Normalize the data

Raw scraped data often contains inconsistent formats.

Examples:

  • Prices like $12.99, 12,99 EUR, USD 12.99
  • Dates like Jan 3, 2025 or 2025-01-03
  • Names with extra whitespace
  • Numbers with commas

Normalize them into standard types:

  • strings -> trimmed
  • prices -> numeric + currency
  • dates -> ISO 8601
  • booleans -> true/false
  • lists -> proper arrays or separate table rows

Example:

import re
from datetime import datetime

def parse_price(text):
    match = re.search(r"[\d,.]+", text)
    return float(match.group().replace(",", "")) if match else None

def parse_date(text):
    return datetime.strptime(text, "%b %d, %Y").date().isoformat()

5) Clean the dataset

Typical cleaning steps:

  • Remove duplicates
  • Handle missing values
  • Standardize categories
  • Filter out bad rows
  • Deduplicate by URL or ID
  • Strip HTML tags and whitespace
  • Fix encoding issues

Example checks:

  • Is price numeric?
  • Is url valid?
  • Is name empty?
  • Are there duplicate records?

6) Validate the output

Don’t assume extraction worked.

Useful validation:

  • Count rows vs expected pages
  • Check null rates per column
  • Verify sample records manually
  • Check outliers
  • Assert schema types

Tools:

  • pandas
  • pandera
  • great_expectations

Example with pandas:

import pandas as pd

df = pd.read_json("data.jsonl", lines=True)
print(df.isna().mean())
print(df.head())

7) Store it in a usable format

Good options:

  • CSV: simple, human-readable
  • JSONL: great for nested or row-based records
  • Parquet: best for analytics and large datasets
  • SQL database: good for querying and updates

Recommended:

  • Use JSONL during scraping
  • Convert to Parquet or a database after cleaning

8) Handle difficult cases

Dynamic pages

Use a browser automation tool:

  • Playwright is often the best choice today.

Sites with anti-bot measures

  • Slow down requests
  • Rotate user agents responsibly
  • Use browser rendering if needed
  • Avoid anything that violates site rules

Mixed layouts

Use:

  • Page-type classifiers
  • Multiple extraction rules
  • LLM fallback for edge cases

Noisy text extraction

Try boilerplate removal libraries or main-content extractors.


9) A simple practical pipeline

A common workflow:

  1. Crawl URLs
  2. Download HTML
  3. Extract fields with selectors or JSON-LD
  4. Normalize and clean fields
  5. Validate schema
  6. Export to CSV/Parquet

10) Example end-to-end structure

For a scraping project, you might organize it like this:

project/
  crawl.py
  extract.py
  clean.py
  validate.py
  data/
    raw/
    processed/
  • crawl.py saves raw HTML
  • extract.py parses HTML into records
  • clean.py standardizes values
  • validate.py checks quality

When to use an LLM

Use an LLM when:

  • Page structures vary a lot
  • You need to extract semantic fields from messy content
  • You have moderate data volume

Avoid relying on it alone when:

  • You need strict reproducibility
  • You’re processing millions of pages
  • You need guaranteed exact parsing

Best pattern:

  • Deterministic parsing first
  • LLM only as fallback

If you want, I can also show you:

  1. a Python scraping template,
  2. a Playwright example for JS-heavy pages, or
  3. a pipeline that uses an LLM to extract structured JSON from HTML.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.