Prompt

What are the most common reasons a web scraper stops working?

Technology / Data Infrastructure3 observationsLast seen Aug 27, 2026

Latest observation

Aug 27, 2026Gemini APIWeb search: off

Web scrapers are notoriously fragile because they rely on the structure, performance, and security policies of external websites—none of which you control.

Here are the most common reasons web scrapers stop working, categorized by the root cause:


1. Changes to the Website’s HTML/DOM Structure

This is the #1 reason scrapers break. Developers update websites constantly, which alters the underlying HTML code.

  • Changed Selectors: A class name like .product-price might be renamed to .price-val or replaced entirely.
  • Dynamic/Hashed Class Names: Modern web frameworks (like React or Tailwind) often generate randomized class names (e.g., <div class="sc-bdVaJa eKjdZs">) that change every time the website is rebuilt.
  • Restructured Layout: Moving an element from inside a <div> to a <section> breaks hardcoded XPath selectors.

2. Anti-Bot and Security Countermeasures

Websites actively try to block automated traffic to protect their data and server load.

  • IP Rate Limiting & Bans: If your scraper sends too many requests per second from a single IP address, the server will block that IP (returning HTTP 429 Too Many Requests or 403 Forbidden).
  • Cloudflare / Web Application Firewalls (WAFs): Intermediary security services (Cloudflare, Akamai, PerimeterX) intercept requests and block scraper traffic before it even reaches the site.
  • CAPTCHAs: Sites automatically trigger reCAPTCHA, HCaptcha, or Cloudflare Turnstile when suspicious automated behavior is detected.
  • Browser Fingerprinting: Anti-bot systems inspect TLS fingerprints, HTTP headers, canvas rendering, and browser flags (e.g., detecting navigator.webdriver = true in Selenium/Puppeteer).

3. Shift to Dynamic JavaScript Rendering

  • Single Page Applications (SPAs): If a website updates its tech stack to React, Angular, or Vue, the raw HTML returned by simple HTTP requests (like Python's requests library) might be completely empty, with content rendered via JavaScript later.
  • Lazy Loading / Infinite Scroll: Content may only load when a user scrolls down or interacts with the page, breaking scrapers that only fetch the initial HTML payload.
  • Hidden API Changes: If your scraper relies on sniffing internal JSON API endpoints instead of parsing HTML, developers frequently change API URL parameters, headers, or authentication payloads.

4. Authentication and Session Expiration

If your scraper logs into a website to fetch data behind a paywall or user profile:

  • Expired Cookies/Tokens: Auth tokens (like JWTs) or session cookies expire after a set time.
  • CSRF/XSRF Tokens: Many sites require dynamic Cross-Site Request Forgery tokens that must be extracted from the page before submitting a form or payload.
  • Forced Logouts: Frequent requests from unexpected locations/devices can trigger mandatory password resets or 2FA prompts.

5. Geo-Targeting and Localization (A/B Testing)

  • Location-Based Content: A website might render entirely different HTML, currency, or product lists based on the IP address location of your scraper (or proxy).
  • A/B Testing: E-commerce sites constantly test different layouts on different users. Your scraper might get Layout A 50% of the time and Layout B the other 50%, causing intermittent failures.

6. Network and Infrastructure Failures

  • Bad/Dead Proxies: If you use a pool of proxies, individual proxy servers often die, slow down, or get blocked, leading to connection timeouts.
  • Target Server Downtime: The target website itself might experience temporary outages (HTTP 500, 502, or 503 errors).
  • Unresolved Timeouts: Scraping scripts without proper timeout settings can hang indefinitely if a page takes too long to load.

How to Make Scrapers More Resilient:

  1. Use Robust Selectors: Prefer target IDs, data attributes (e.g., data-testid="price"), or relative XPaths over fragile CSS class names.
  2. Scrape APIs Directly: Look at the browser's Network tab. Fetching data directly from internal JSON APIs is much faster and cleaner than parsing HTML.
  3. Rotate IPs and User-Agents: Use residential proxy pools and randomize HTTP headers.
  4. Use Headless Browsers with Stealth Plugins: Tools like Playwright or Puppeteer paired with stealth plugins help bypass basic browser fingerprinting.
  5. Implement Retries and Alerting: Use automatic retry logic with exponential backoff for failed requests, and set up alerts (like Slack/Email notifications) when your scraper's output drops to zero.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (3 observations for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.