Prompt
How to extract product data from multiple domains
Latest observation
Extracting product data from multiple domains is usually done with a combination of site discovery, page fetching, parsing, normalization, and storage. The exact approach depends on whether you’re allowed to scrape the sites and how consistent the target pages are.
1) Start with the source type
Different domains expose product data in different ways:
- Structured feeds: CSV, JSON, XML, RSS
- Public APIs: best option if available
- HTML product pages: most common scraping target
- JavaScript-rendered pages: may need a headless browser
2) Preferred extraction methods
Best case: API or feed
If a site has:
- a product API
- a sitemap
- a Google Merchant feed
- JSON-LD structured data
Use that instead of scraping HTML.
HTML scraping
For sites without APIs:
- fetch the page
- parse product name, price, SKU, description, images, availability
- use CSS selectors/XPath or structured data extraction
JSON-LD / Microdata
Many ecommerce sites embed product info in schema.org markup:
application/ld+jsonProductOfferAggregateRating
This is often more stable than scraping page text.
3) Handle multiple domains with a modular pipeline
A good design is:
A. Domain discovery
Maintain a list of target domains and their product page patterns.
B. Domain-specific extractors
Each domain may need its own parser:
- selectors for title/price
- pagination logic
- category URL patterns
C. Common normalized schema
Map all extracted data into one format, for example:
{
"source_domain": "example.com",
"product_url": "...",
"name": "...",
"sku": "...",
"price": 19.99,
"currency": "USD",
"availability": "in_stock",
"description": "...",
"images": ["..."],
"brand": "..."
}
D. Storage
Store in:
- database (PostgreSQL, MongoDB)
- CSV/Parquet
- search index
- data warehouse
4) Recommended workflow
- Collect URLs
From sitemap, category pages, search results, or crawl links. - Fetch pages
Userequestsfor static pages orPlaywright/Seleniumfor JS sites. - Parse product data
- first try JSON-LD
- fallback to HTML selectors
- Clean and normalize
- currency conversion if needed
- convert prices to numbers
- standardize availability values
- Deduplicate
Use product URL, SKU, or normalized name+brand. - Validate
Ensure required fields exist and prices are sane. - Store and monitor
Re-run regularly and track changes.
5) Tools commonly used
Python
requestsBeautifulSoup4lxmlscrapyplaywrightpandasextructfor structured data extraction
For large-scale crawling
- Scrapy + distributed workers
- Airflow / Prefect for scheduling
- Redis / Kafka for queues
6) Example approach in Python
A simple strategy is:
- fetch HTML
- extract JSON-LD if present
- otherwise parse selectors
import requests
from bs4 import BeautifulSoup
import json
url = "https://example.com/product/123"
html = requests.get(url, timeout=20).text
soup = BeautifulSoup(html, "html.parser")
# Try JSON-LD
product_data = None
for script in soup.find_all("script", type="application/ld+json"):
try:
data = json.loads(script.string)
if isinstance(data, dict) and data.get("@type") == "Product":
product_data = data
break
except Exception:
pass
if product_data:
print(product_data.get("name"))
print(product_data.get("offers", {}).get("price"))
else:
# Fallback selectors
name = soup.select_one("h1.product-title")
price = soup.select_one(".price")
print(name.get_text(strip=True) if name else None)
print(price.get_text(strip=True) if price else None)
7) Challenges with multiple domains
- every site has different HTML structure
- anti-bot protections
- rate limits and CAPTCHAs
- dynamic rendering
- inconsistent product fields
- locale/currency differences
8) Good practices
- check
robots.txtand terms of service - use respectful rate limits
- identify your crawler with a user agent
- cache responses
- log failures per domain
- use retries with backoff
9) If you need many domains at scale
Consider a configuration-driven setup:
domain: example.com
start_urls:
- https://example.com/category/shoes
selectors:
name: "h1.product-title"
price: ".price"
sku: ".sku"
Then write one crawler that reads configs for each domain.
If you want, I can also give you:
- a Python Scrapy project structure for multi-domain product extraction, or
- a custom scraper template that handles JSON-LD + HTML fallback.