Prompt

Why is my crawler getting blocked after a few hundred requests?

Technology / Cloud Infrastructure1 observationLast seen Oct 4, 2026

Latest observation

Oct 4, 2026OpenAI APIWeb search: off

A crawler usually gets blocked after a few hundred requests because the site is detecting traffic that looks automated, abusive, or unusual. Common causes:

Typical reasons

  • Rate limiting: too many requests from one IP in a short time.
  • Bot detection: repeated patterns, missing browser headers, no JS execution, etc.
  • IP reputation: datacenter/VPN/proxy IPs may be flagged more easily.
  • Lack of session behavior: no cookies, no referrers, no navigation flow.
  • Too many concurrent requests: bursts can trigger defenses.
  • Duplicate/fingerprintable behavior: same user-agent, same timings, same paths.
  • Missing authentication or CSRF tokens where the site expects them.
  • Robots / access policy enforcement: some sites actively block scraping.

Things to check

  • Compare status codes: 403, 429, 503, CAPTCHA pages, redirects to challenge pages.
  • Inspect the response body: sometimes it’s a block page disguised as HTML.
  • Review request headers and cookies: are you sending what a real browser would?
  • Measure request rate and concurrency: do blocks happen at a predictable threshold?
  • Check whether the site uses a WAF/CDN like Cloudflare, Akamai, or Datadome.

What usually helps

  • Slow down and add jitter to delays.
  • Reduce parallelism.
  • Respect robots.txt and site terms.
  • Use conditional requests and caching to avoid re-fetching.
  • Keep a more realistic session/cookie flow if allowed.
  • Handle 429/403 gracefully and back off.

Important note

If the site is intentionally blocking automation, the right approach is usually to obtain permission, use an official API, or limit crawling to what’s permitted. I can help you diagnose the likely cause if you share:

  • the HTTP status codes you’re seeing,
  • the request rate/concurrency,
  • and a sample of the response headers/body from a blocked request.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.