Prompt

How to extract product data from multiple domains

Technology / Cloud Infrastructure1 observationLast seen Oct 4, 2026

Latest observation

Oct 4, 2026OpenAI APIWeb search: off

Extracting product data from multiple domains is usually done with a combination of site discovery, page fetching, parsing, normalization, and storage. The exact approach depends on whether you’re allowed to scrape the sites and how consistent the target pages are.

1) Start with the source type

Different domains expose product data in different ways:

  • Structured feeds: CSV, JSON, XML, RSS
  • Public APIs: best option if available
  • HTML product pages: most common scraping target
  • JavaScript-rendered pages: may need a headless browser

2) Preferred extraction methods

Best case: API or feed

If a site has:

  • a product API
  • a sitemap
  • a Google Merchant feed
  • JSON-LD structured data

Use that instead of scraping HTML.

HTML scraping

For sites without APIs:

  • fetch the page
  • parse product name, price, SKU, description, images, availability
  • use CSS selectors/XPath or structured data extraction

JSON-LD / Microdata

Many ecommerce sites embed product info in schema.org markup:

  • application/ld+json
  • Product
  • Offer
  • AggregateRating

This is often more stable than scraping page text.

3) Handle multiple domains with a modular pipeline

A good design is:

A. Domain discovery

Maintain a list of target domains and their product page patterns.

B. Domain-specific extractors

Each domain may need its own parser:

  • selectors for title/price
  • pagination logic
  • category URL patterns

C. Common normalized schema

Map all extracted data into one format, for example:

{
  "source_domain": "example.com",
  "product_url": "...",
  "name": "...",
  "sku": "...",
  "price": 19.99,
  "currency": "USD",
  "availability": "in_stock",
  "description": "...",
  "images": ["..."],
  "brand": "..."
}

D. Storage

Store in:

  • database (PostgreSQL, MongoDB)
  • CSV/Parquet
  • search index
  • data warehouse

4) Recommended workflow

  1. Collect URLs
    From sitemap, category pages, search results, or crawl links.
  2. Fetch pages
    Use requests for static pages or Playwright/Selenium for JS sites.
  3. Parse product data
    • first try JSON-LD
    • fallback to HTML selectors
  4. Clean and normalize
    • currency conversion if needed
    • convert prices to numbers
    • standardize availability values
  5. Deduplicate
    Use product URL, SKU, or normalized name+brand.
  6. Validate
    Ensure required fields exist and prices are sane.
  7. Store and monitor
    Re-run regularly and track changes.

5) Tools commonly used

Python

  • requests
  • BeautifulSoup4
  • lxml
  • scrapy
  • playwright
  • pandas
  • extruct for structured data extraction

For large-scale crawling

  • Scrapy + distributed workers
  • Airflow / Prefect for scheduling
  • Redis / Kafka for queues

6) Example approach in Python

A simple strategy is:

  • fetch HTML
  • extract JSON-LD if present
  • otherwise parse selectors
import requests
from bs4 import BeautifulSoup
import json

url = "https://example.com/product/123"
html = requests.get(url, timeout=20).text
soup = BeautifulSoup(html, "html.parser")

# Try JSON-LD
product_data = None
for script in soup.find_all("script", type="application/ld+json"):
    try:
        data = json.loads(script.string)
        if isinstance(data, dict) and data.get("@type") == "Product":
            product_data = data
            break
    except Exception:
        pass

if product_data:
    print(product_data.get("name"))
    print(product_data.get("offers", {}).get("price"))
else:
    # Fallback selectors
    name = soup.select_one("h1.product-title")
    price = soup.select_one(".price")
    print(name.get_text(strip=True) if name else None)
    print(price.get_text(strip=True) if price else None)

7) Challenges with multiple domains

  • every site has different HTML structure
  • anti-bot protections
  • rate limits and CAPTCHAs
  • dynamic rendering
  • inconsistent product fields
  • locale/currency differences

8) Good practices

  • check robots.txt and terms of service
  • use respectful rate limits
  • identify your crawler with a user agent
  • cache responses
  • log failures per domain
  • use retries with backoff

9) If you need many domains at scale

Consider a configuration-driven setup:

domain: example.com
start_urls:
  - https://example.com/category/shoes
selectors:
  name: "h1.product-title"
  price: ".price"
  sku: ".sku"

Then write one crawler that reads configs for each domain.


If you want, I can also give you:

  1. a Python Scrapy project structure for multi-domain product extraction, or
  2. a custom scraper template that handles JSON-LD + HTML fallback.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as Circuit. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.