Prompt
How do I crawl product pages across many ecommerce sites?
Latest observation
Crawling product pages across many ecommerce sites is mostly a systems + compliance problem, not just a scraping problem. The safest approach is to combine discovery, fetching, extraction, politeness, and site-specific adaptation.
1) Start with the legal / access side
Before crawling:
- Check each site’s Terms of Service
- Review robots.txt
- Respect rate limits and avoid bypassing access controls
- Prefer publicly available product data and official APIs or feeds when available
- If you need data at scale, consider getting permission or using a data provider
2) Build a crawler that discovers product URLs
Common discovery methods:
- Sitemaps: look for
/sitemap.xmland nested sitemap indexes - Category/listing pages: crawl category pages and pagination
- Internal links: follow links that look like product URLs
- Structured data: product pages often contain
schema.org/Product - Search endpoints: only if allowed and documented
A good strategy is:
- Seed with homepage/category/sitemap URLs
- Filter for likely product URLs
- Queue and crawl only those URLs
- Extract product data from the page
3) Detect product pages reliably
Use multiple signals:
- URL patterns:
/product/,/p/,/item/ - HTML metadata:
<meta property="og:type" content="product">- JSON-LD with
@type: Product
- Page features:
- price
- availability
- add-to-cart button
- product title
Don’t rely on URL patterns alone; many sites use different structures.
4) Prefer structured data extraction first
Most ecommerce sites expose useful data in:
- JSON-LD
- Open Graph
- Microdata / RDFa
- Embedded state blobs in scripts
Example JSON-LD fields to extract:
- name
- description
- sku
- brand
- offers.price
- offers.priceCurrency
- offers.availability
- aggregateRating
- image
This is usually more stable than scraping rendered HTML text.
5) Handle site differences with a layered extractor
A practical architecture:
- Generic extractor for common ecommerce markup
- Site adapters for exceptions
- Fallback HTML parsing when structured data is missing
- Headless browser rendering only for pages that require JavaScript
This lets you scale without writing custom code for every site.
6) Use polite crawling practices
To avoid overwhelming sites:
- Set a reasonable concurrency per domain
- Add delays between requests
- Respect
Retry-After - Cache results and use ETag / If-Modified-Since if available
- Identify your crawler in the User-Agent
- Stop on repeated errors or blocks
7) Watch for anti-bot and dynamic content
Many ecommerce sites use:
- JavaScript-rendered pages
- lazy-loaded data
- bot protection / rate limiting
- A/B testing
Use:
- plain HTTP fetching first
- headless browser only when needed
- session reuse carefully
- robust retry/backoff logic
Avoid trying to bypass bot protections; if access is restricted, get authorization or use another source.
8) Normalize the product data
Across sites, map fields into a standard schema:
product_idurltitlebrandcategorypricecurrencyavailabilityimagesdescriptionvariantsshippingrating
This makes downstream analysis much easier.
9) Store crawl state and deduplicate
You’ll want:
- a URL frontier / queue
- visited URL set
- canonical URL handling
- deduplication by product identifiers or normalized titles
- change tracking for price/availability updates
10) A practical stack
Common options:
- Scrapy for crawling
- Playwright or Selenium for JS-rendered pages
- BeautifulSoup/lxml for parsing
- Redis / Kafka / PostgreSQL for queues and storage
- Pydantic or similar for structured product records
11) Suggested workflow
- Collect sitemap and category seeds for each domain
- Crawl with a polite scheduler
- Detect product pages with rules + structured data
- Extract data using JSON-LD first
- Use browser rendering only when necessary
- Normalize into a common schema
- Monitor failures and add site-specific adapters over time
12) Minimal implementation idea
At a high level:
- Discovery
- fetch sitemap/category pages
- collect candidate URLs
- Classification
- decide if URL is a product page
- Extraction
- parse JSON-LD / meta tags / HTML
- Politeness
- per-domain rate limiting
- Storage
- save normalized records
If you want, I can help you with one of these next:
- a Scrapy-based crawler architecture
- a Python script to extract product data from JSON-LD
- a site-agnostic product page detector
- a full example pipeline for multiple ecommerce domains