Prompt
How do I pull competitor pricing from websites reliably when pages keep changing?
Latest observation
Pulling competitor pricing reliably is mostly a data engineering + change management problem, not just a scraping problem. The key is to design for variation, detection, and fallback.
1) Prefer structured sources first
Before scraping pages, look for:
- Public APIs
- Product feeds / JSON-LD / schema.org markup
- Sitemaps
- Price list PDFs or downloadable catalogs
- Embedded JSON in page source (
application/ld+json,__NEXT_DATA__,window.__INITIAL_STATE__, etc.)
These are usually more stable than visible HTML.
2) Build scrapers that don’t depend on exact page layout
Avoid brittle selectors like:
div:nth-child(3) > span.price
Instead use:
- Semantic anchors: product name, SKU, “price”, currency symbols
- Multiple fallback selectors
- Regex extraction for price patterns
- DOM traversal relative to stable labels
Example strategy:
- Try structured data
- Try known JSON state blobs
- Try page text + regex
- Try HTML selectors as last resort
3) Normalize the data model
Store prices in a consistent schema:
competitorproduct_id/skuproduct_namepricecurrencyunit(if applicable)promo_pricelist_priceavailabilitycaptured_atsource_urlconfidence_score
This makes changes easier to detect and compare.
4) Use change detection and validation
Pages will change, so detect when extraction becomes suspicious:
- Price missing or zero
- Price format changed
- Large unexpected jumps/drops
- Product name mismatch
- Currency mismatch
- Too many pages suddenly failing
Add alerts for:
- Selector failure rate spike
- Parsing confidence drop
- HTML template hash changes
- Output anomalies compared to historical data
5) Store raw snapshots
Save:
- Raw HTML
- Extracted JSON
- Screenshot if needed
This lets you debug failures and reprocess old pages when your parser improves.
6) Make parsing modular
Use a per-site parser architecture:
- One common crawler
- Site-specific extraction plugins
- Shared normalization and validation layer
Example:
fetcherrenderer(if JS needed)extractor_site_aextractor_site_bvalidatorstore
7) Handle JavaScript-heavy sites
If pricing is rendered client-side:
- Use a headless browser only when necessary
- Wait for network idle or a specific selector
- Prefer intercepting XHR/fetch calls that return pricing data
- Cache rendered output to reduce load
8) Track template versions
For each site, compute and compare:
- HTML structure fingerprints
- Key DOM node paths
- Script bundle versions
- Presence/absence of known markers
When template changes, route to a review queue rather than silently failing.
9) Add human-in-the-loop review for exceptions
For high-value competitors:
- Surface failures in a dashboard
- Let an analyst confirm corrected selectors
- Promote fixes into the site-specific parser
10) Respect legal and technical constraints
Make sure you:
- Follow robots.txt and site terms where applicable
- Don’t bypass authentication, paywalls, or anti-bot protections improperly
- Rate-limit requests and identify yourself if appropriate
- Prefer official partnerships or feeds for scale and reliability
A practical reliability pattern
A strong pipeline looks like this:
- Discover URLs
- Fetch page
- Check for structured data
- Extract using site-specific rules
- Validate output
- Compare against historical expectations
- Store result + raw snapshot
- Alert on anomalies
- Review failures and update parser
Useful signals to monitor
- Extraction success rate
- Median price change by site/product
- Selector hit rate
- JS rendering dependency rate
- Time-to-refresh after site changes
- % of pages requiring manual review
If you want the simplest reliable approach
For each competitor site:
- Start with JSON-LD / embedded data
- Write a site-specific parser with 2–3 fallback methods
- Save raw HTML and extracted values
- Add anomaly detection
- Maintain a small review loop
If you want, I can also help you design a robust competitor pricing scraper architecture or give you a sample Python implementation with fallbacks and change detection.