Prompt
How do I scrape product prices from websites at scale?
Latest observation
Scraping product prices at scale is usually a data engineering + anti-bot + compliance problem, not just a “write a scraper” problem.
1) First, use the cheapest legitimate source
Before scraping, check for:
- Official APIs from the retailer/marketplace
- Affiliate / partner feeds
- XML/CSV product feeds
- Sitemaps or structured data (
schema.org/Product, JSON-LD) - Internal pages or endpoints already used by the site (only if allowed by the site’s terms)
These are more stable and far easier to scale than browser scraping.
2) If you must scrape, build it like a pipeline
A scalable price-collection system usually has these parts:
A. URL discovery
- Maintain a catalog of product URLs
- Track canonical URLs and variants
- Refresh only products that need updates
B. Fetching layer
Use an HTTP client first, browser automation only when necessary:
- HTTP scraping for static pages
- Headless browser for JS-rendered pages
At scale:
- Use a queue of URLs
- Run workers in parallel
- Add rate limiting per domain
- Use retries with exponential backoff
- Cache responses where possible
C. Parsing layer
Extract:
- product name
- current price
- currency
- availability
- variant info
- timestamp
- source URL
Prefer:
- JSON-LD
- embedded data blobs
- structured HTML selectors
- site-specific parsers per domain
D. Storage and change tracking
Store every observation with:
- product ID
- price
- timestamp
- source
- confidence/status
Then compute:
- price history
- alerts
- deltas
- min/max over time
3) Handle anti-bot measures carefully
Many sites actively block automated traffic. Legitimate scaling requires respecting site policies and avoiding disruptive behavior.
Practical considerations:
- Identify yourself with a clear User-Agent if appropriate
- Keep request rates low and consistent
- Avoid hammering pages repeatedly
- Respect robots.txt where applicable
- Don’t bypass login, CAPTCHAs, or access controls
- If the site blocks you, prefer an official feed/API or permission
4) Recommended architecture
A common setup:
- Scheduler: decides what to refresh
- Queue: RabbitMQ / SQS / Redis queue / Kafka
- Workers: fetch pages, extract price
- Parser service: site-specific extraction logic
- DB: PostgreSQL / BigQuery / ClickHouse
- Monitoring: failures, bans, parse drift, latency
- Alerting: when prices change or extraction breaks
5) Tech stack suggestions
For smaller scale
- Python
- Requests / httpx
- BeautifulSoup / lxml
- PostgreSQL
- Celery or RQ
For larger scale
- Python or Go for fetchers
- Scrapy for crawling orchestration
- Playwright for JS-heavy pages
- Redis/SQS/Kafka for queues
- Docker + Kubernetes for worker autoscaling
- ClickHouse/BigQuery for analytics
6) Make scrapers resilient
Websites change often. Add:
- Selector fallbacks
- Schema-based extraction first, HTML fallback second
- Validation rules:
- price must be numeric
- currency must be known
- reject absurd values
- Per-site tests with sample pages
- Monitoring for parse failures
7) Respect legal and operational constraints
At scale, the biggest risks are:
- violating terms of service
- collecting data you’re not allowed to collect
- overloading sites
- inaccurate or stale data
- expensive maintenance
If your use case is commercial, it’s worth getting:
- written permission
- a data license
- or a partner feed
8) A practical workflow
- Start with 5–10 target sites
- For each, inspect page source for structured price data
- Build one parser per site
- Add a queue and worker pool
- Track extraction success rate and parse drift
- Expand only after reliability is good
- Replace scraping with APIs/feeds where possible
9) Minimal example approach
A simple pattern:
- Fetch page
- Extract JSON-LD or meta tags
- Fallback to CSS selectors
- Validate price
- Save snapshot
If you want, I can help you with one of these next:
- a Python scraper template
- a Scrapy + Playwright architecture
- a price extraction strategy for a specific site
- a database schema for price history
- a scalable queue/worker design