Prompt

How do I pull competitor pricing from websites reliably when pages keep changing?

Technology / Data infrastructure1 observationLast seen Sep 24, 2026

Latest observation

Sep 24, 2026OpenAI APIWeb search: off

Pulling competitor pricing reliably is mostly a data engineering + change management problem, not just a scraping problem. The key is to design for variation, detection, and fallback.

1) Prefer structured sources first

Before scraping pages, look for:

  • Public APIs
  • Product feeds / JSON-LD / schema.org markup
  • Sitemaps
  • Price list PDFs or downloadable catalogs
  • Embedded JSON in page source (application/ld+json, __NEXT_DATA__, window.__INITIAL_STATE__, etc.)

These are usually more stable than visible HTML.

2) Build scrapers that don’t depend on exact page layout

Avoid brittle selectors like:

  • div:nth-child(3) > span.price

Instead use:

  • Semantic anchors: product name, SKU, “price”, currency symbols
  • Multiple fallback selectors
  • Regex extraction for price patterns
  • DOM traversal relative to stable labels

Example strategy:

  1. Try structured data
  2. Try known JSON state blobs
  3. Try page text + regex
  4. Try HTML selectors as last resort

3) Normalize the data model

Store prices in a consistent schema:

  • competitor
  • product_id / sku
  • product_name
  • price
  • currency
  • unit (if applicable)
  • promo_price
  • list_price
  • availability
  • captured_at
  • source_url
  • confidence_score

This makes changes easier to detect and compare.

4) Use change detection and validation

Pages will change, so detect when extraction becomes suspicious:

  • Price missing or zero
  • Price format changed
  • Large unexpected jumps/drops
  • Product name mismatch
  • Currency mismatch
  • Too many pages suddenly failing

Add alerts for:

  • Selector failure rate spike
  • Parsing confidence drop
  • HTML template hash changes
  • Output anomalies compared to historical data

5) Store raw snapshots

Save:

  • Raw HTML
  • Extracted JSON
  • Screenshot if needed

This lets you debug failures and reprocess old pages when your parser improves.

6) Make parsing modular

Use a per-site parser architecture:

  • One common crawler
  • Site-specific extraction plugins
  • Shared normalization and validation layer

Example:

  • fetcher
  • renderer (if JS needed)
  • extractor_site_a
  • extractor_site_b
  • validator
  • store

7) Handle JavaScript-heavy sites

If pricing is rendered client-side:

  • Use a headless browser only when necessary
  • Wait for network idle or a specific selector
  • Prefer intercepting XHR/fetch calls that return pricing data
  • Cache rendered output to reduce load

8) Track template versions

For each site, compute and compare:

  • HTML structure fingerprints
  • Key DOM node paths
  • Script bundle versions
  • Presence/absence of known markers

When template changes, route to a review queue rather than silently failing.

9) Add human-in-the-loop review for exceptions

For high-value competitors:

  • Surface failures in a dashboard
  • Let an analyst confirm corrected selectors
  • Promote fixes into the site-specific parser

10) Respect legal and technical constraints

Make sure you:

  • Follow robots.txt and site terms where applicable
  • Don’t bypass authentication, paywalls, or anti-bot protections improperly
  • Rate-limit requests and identify yourself if appropriate
  • Prefer official partnerships or feeds for scale and reliability

A practical reliability pattern

A strong pipeline looks like this:

  1. Discover URLs
  2. Fetch page
  3. Check for structured data
  4. Extract using site-specific rules
  5. Validate output
  6. Compare against historical expectations
  7. Store result + raw snapshot
  8. Alert on anomalies
  9. Review failures and update parser

Useful signals to monitor

  • Extraction success rate
  • Median price change by site/product
  • Selector hit rate
  • JS rendering dependency rate
  • Time-to-refresh after site changes
  • % of pages requiring manual review

If you want the simplest reliable approach

For each competitor site:

  • Start with JSON-LD / embedded data
  • Write a site-specific parser with 2–3 fallback methods
  • Save raw HTML and extracted values
  • Add anomaly detection
  • Maintain a small review loop

If you want, I can also help you design a robust competitor pricing scraper architecture or give you a sample Python implementation with fallbacks and change detection.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.