Prompt

How do I create a reliable scraper for rapidly changing websites?

Technology / Data infrastructure1 observationLast seen Sep 24, 2026

Latest observation

Sep 24, 2026OpenAI APIWeb search: off

Creating a reliable scraper for a rapidly changing website is mostly about resilience, observability, and minimizing coupling to fragile page structure. Here’s a practical approach.

1) Prefer stable data sources over HTML scraping

Before scraping the rendered page, check for:

  • Official APIs
  • Hidden JSON endpoints used by the frontend
  • GraphQL endpoints
  • Network requests in browser dev tools

These are usually more stable than page markup and easier to parse.

2) Use selectors that are resilient

Avoid brittle selectors like long absolute paths or dynamic class names.

Better options:

  • Semantic attributes like data-*
  • Text anchors near the target element
  • ARIA labels
  • Stable IDs
  • XPath/CSS selectors based on structure, not volatile styling

Example:

  • Fragile: .container > div:nth-child(3) > span.price
  • Better: [data-testid="price"] or span[aria-label="price"]

3) Build in multiple extraction strategies

Assume the site will change.

For each field, define fallback logic:

  1. Primary selector
  2. Secondary selector
  3. Extract from embedded JSON
  4. Extract from neighboring context
  5. Mark as missing and continue

This prevents one DOM change from breaking the whole scraper.

4) Separate fetching, parsing, and storage

Use a pipeline like:

  • Fetcher: downloads HTML / JSON
  • Parser: extracts fields
  • Validator: checks completeness and format
  • Writer: stores results

That makes debugging much easier when something breaks.

5) Add validation and anomaly detection

Detect when the site changes by checking:

  • Missing fields
  • Sudden drop in item count
  • Unexpected HTML structure
  • Changed content type
  • Parsing errors
  • Unusual response codes or redirects

Example alerts:

  • “Parsed 0 products instead of usual 50”
  • “Price field missing in 80% of records”
  • “Response now contains login page”

6) Log raw snapshots for failed cases

When parsing fails, save:

  • The raw HTML or JSON
  • URL
  • Timestamp
  • Response headers
  • Selector used
  • Error details

This makes fixing the scraper much faster.

7) Use rate limiting and retries carefully

Rapid changes often come with anti-bot measures or unstable responses.

Use:

  • Exponential backoff
  • Retry on transient failures
  • Randomized delay/jitter
  • Concurrency limits
  • Respect for robots.txt and site terms where applicable

Avoid aggressive retries that amplify blocking.

8) Handle dynamic content explicitly

If content is loaded by JavaScript:

  • Use a browser automation tool only when necessary
  • Wait for specific elements or network calls, not arbitrary sleep
  • Prefer API calls discovered from the page over full browser rendering when possible

Tools:

  • Playwright
  • Selenium
  • Puppeteer

9) Make selectors and field mappings configurable

Don’t hardcode everything in code. Put selectors in config files so you can update them quickly without redeploying the whole scraper.

Example config:

{
  "product_name": ["[data-testid='name']", "h1.product-title"],
  "price": ["[data-testid='price']", ".price"]
}

10) Version your parsers

When the website changes:

  • Keep old parser versions
  • Compare outputs between versions
  • Roll back quickly if a fix causes regressions

11) Test against saved HTML samples

Maintain a fixture set of:

  • Normal pages
  • Edge cases
  • Missing fields
  • Updated versions after site changes

Run tests whenever you update parsing logic.

12) Use monitoring and alerts

Track:

  • Success rate
  • Field coverage
  • Parse latency
  • Block rate
  • Response size
  • Item count trends

This helps you detect failure before downstream data consumers do.


A practical architecture

A robust scraper often looks like this:

  1. Discover URLs
  2. Fetch page or API response
  3. Detect page type/version
  4. Parse with layered fallback extractors
  5. Validate fields
  6. Store result
  7. Log failures with raw input
  8. Alert on anomalies

If the site changes very frequently

Consider these additional tactics:

  • Scrape from backend endpoints instead of frontend HTML
  • Use machine-readable structured data if available (JSON-LD, Open Graph, schema.org)
  • Monitor the site’s network requests periodically
  • Keep parsers small and modular
  • Use diffing to compare new page versions with old ones

Important legal/ethical note

Make sure your scraping complies with:

  • The site’s terms of service
  • Robots.txt where relevant
  • Copyright/privacy rules
  • Rate limits and access policies

If you want, I can also show you:

  1. a Python scraper template with fallback selectors, or
  2. a Playwright-based scraper that handles dynamic pages reliably.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as Circuit. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.