Prompt

structured data extraction from web pages

Technology / Cloud Infrastructure1 observationLast seen Oct 4, 2026

Latest observation

Oct 4, 2026OpenAI APIWeb search: off

Structured data extraction from web pages means turning information on a webpage into organized fields you can store, search, analyze, or feed into another system.

Common approaches

  1. Direct HTML parsing

    • Extract content from the DOM using CSS selectors or XPath.
    • Best when page structure is stable.
  2. Embedded structured data

    • Many pages include:
      • JSON-LD
      • Microdata
      • RDFa
    • Often the cleanest source for things like products, articles, events, recipes.
  3. Browser automation / rendering

    • Use Playwright, Puppeteer, Selenium, etc. to handle JavaScript-rendered pages.
    • Needed when content is loaded dynamically.
  4. Text + LLM-based extraction

    • Convert page content to text and ask a model to extract fields into a schema.
    • Useful for messy or inconsistent pages, but less deterministic.

Typical pipeline

  1. Fetch page
  2. Render if needed
  3. Detect and parse structured data
  4. Fall back to DOM extraction
  5. Normalize fields
  6. Validate against a schema
  7. Store in JSON / database / CSV

Example schema

{
  "title": "",
  "author": "",
  "published_date": "",
  "price": "",
  "currency": "",
  "availability": "",
  "url": ""
}

Tools often used

  • Python: BeautifulSoup, lxml, requests, Scrapy, Playwright
  • JavaScript/TypeScript: Cheerio, Playwright, Puppeteer
  • Structured data parsers: extruct, schema.org parsers
  • Validation: Pydantic, JSON Schema

Example: extracting JSON-LD in Python

import requests
from bs4 import BeautifulSoup
import json

html = requests.get("https://example.com").text
soup = BeautifulSoup(html, "html.parser")

for script in soup.find_all("script", type="application/ld+json"):
    try:
        data = json.loads(script.string)
        print(data)
    except Exception:
        pass

Best practices

  • Prefer structured data already present on the page
  • Validate extracted values
  • Handle pagination and duplicate records
  • Respect robots.txt and site terms
  • Add retries, timeouts, and rate limiting
  • Log extraction failures for debugging

If you want, I can also provide:

  • a complete Python scraper template
  • a Playwright-based extractor
  • or help design an extraction schema for your use case.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.