Prompt

How do I crawl product pages across many ecommerce sites?

Technology / Cloud Infrastructure1 observationLast seen Oct 4, 2026

Latest observation

Oct 4, 2026OpenAI APIWeb search: off

Crawling product pages across many ecommerce sites is mostly a systems + compliance problem, not just a scraping problem. The safest approach is to combine discovery, fetching, extraction, politeness, and site-specific adaptation.

1) Start with the legal / access side

Before crawling:

  • Check each site’s Terms of Service
  • Review robots.txt
  • Respect rate limits and avoid bypassing access controls
  • Prefer publicly available product data and official APIs or feeds when available
  • If you need data at scale, consider getting permission or using a data provider

2) Build a crawler that discovers product URLs

Common discovery methods:

  • Sitemaps: look for /sitemap.xml and nested sitemap indexes
  • Category/listing pages: crawl category pages and pagination
  • Internal links: follow links that look like product URLs
  • Structured data: product pages often contain schema.org/Product
  • Search endpoints: only if allowed and documented

A good strategy is:

  1. Seed with homepage/category/sitemap URLs
  2. Filter for likely product URLs
  3. Queue and crawl only those URLs
  4. Extract product data from the page

3) Detect product pages reliably

Use multiple signals:

  • URL patterns: /product/, /p/, /item/
  • HTML metadata:
    • <meta property="og:type" content="product">
    • JSON-LD with @type: Product
  • Page features:
    • price
    • availability
    • add-to-cart button
    • product title

Don’t rely on URL patterns alone; many sites use different structures.

4) Prefer structured data extraction first

Most ecommerce sites expose useful data in:

  • JSON-LD
  • Open Graph
  • Microdata / RDFa
  • Embedded state blobs in scripts

Example JSON-LD fields to extract:

  • name
  • description
  • sku
  • brand
  • offers.price
  • offers.priceCurrency
  • offers.availability
  • aggregateRating
  • image

This is usually more stable than scraping rendered HTML text.

5) Handle site differences with a layered extractor

A practical architecture:

  • Generic extractor for common ecommerce markup
  • Site adapters for exceptions
  • Fallback HTML parsing when structured data is missing
  • Headless browser rendering only for pages that require JavaScript

This lets you scale without writing custom code for every site.

6) Use polite crawling practices

To avoid overwhelming sites:

  • Set a reasonable concurrency per domain
  • Add delays between requests
  • Respect Retry-After
  • Cache results and use ETag / If-Modified-Since if available
  • Identify your crawler in the User-Agent
  • Stop on repeated errors or blocks

7) Watch for anti-bot and dynamic content

Many ecommerce sites use:

  • JavaScript-rendered pages
  • lazy-loaded data
  • bot protection / rate limiting
  • A/B testing

Use:

  • plain HTTP fetching first
  • headless browser only when needed
  • session reuse carefully
  • robust retry/backoff logic

Avoid trying to bypass bot protections; if access is restricted, get authorization or use another source.

8) Normalize the product data

Across sites, map fields into a standard schema:

  • product_id
  • url
  • title
  • brand
  • category
  • price
  • currency
  • availability
  • images
  • description
  • variants
  • shipping
  • rating

This makes downstream analysis much easier.

9) Store crawl state and deduplicate

You’ll want:

  • a URL frontier / queue
  • visited URL set
  • canonical URL handling
  • deduplication by product identifiers or normalized titles
  • change tracking for price/availability updates

10) A practical stack

Common options:

  • Scrapy for crawling
  • Playwright or Selenium for JS-rendered pages
  • BeautifulSoup/lxml for parsing
  • Redis / Kafka / PostgreSQL for queues and storage
  • Pydantic or similar for structured product records

11) Suggested workflow

  1. Collect sitemap and category seeds for each domain
  2. Crawl with a polite scheduler
  3. Detect product pages with rules + structured data
  4. Extract data using JSON-LD first
  5. Use browser rendering only when necessary
  6. Normalize into a common schema
  7. Monitor failures and add site-specific adapters over time

12) Minimal implementation idea

At a high level:

  • Discovery
    • fetch sitemap/category pages
    • collect candidate URLs
  • Classification
    • decide if URL is a product page
  • Extraction
    • parse JSON-LD / meta tags / HTML
  • Politeness
    • per-domain rate limiting
  • Storage
    • save normalized records

If you want, I can help you with one of these next:

  1. a Scrapy-based crawler architecture
  2. a Python script to extract product data from JSON-LD
  3. a site-agnostic product page detector
  4. a full example pipeline for multiple ecommerce domains

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.