Prompt
Why is my crawler getting blocked after a few hundred requests?
Latest observation
Oct 4, 2026
OpenAI APIWeb search: off
A crawler usually gets blocked after a few hundred requests because the site is detecting traffic that looks automated, abusive, or unusual. Common causes:
Typical reasons
- Rate limiting: too many requests from one IP in a short time.
- Bot detection: repeated patterns, missing browser headers, no JS execution, etc.
- IP reputation: datacenter/VPN/proxy IPs may be flagged more easily.
- Lack of session behavior: no cookies, no referrers, no navigation flow.
- Too many concurrent requests: bursts can trigger defenses.
- Duplicate/fingerprintable behavior: same user-agent, same timings, same paths.
- Missing authentication or CSRF tokens where the site expects them.
- Robots / access policy enforcement: some sites actively block scraping.
Things to check
- Compare status codes: 403, 429, 503, CAPTCHA pages, redirects to challenge pages.
- Inspect the response body: sometimes it’s a block page disguised as HTML.
- Review request headers and cookies: are you sending what a real browser would?
- Measure request rate and concurrency: do blocks happen at a predictable threshold?
- Check whether the site uses a WAF/CDN like Cloudflare, Akamai, or Datadome.
What usually helps
- Slow down and add jitter to delays.
- Reduce parallelism.
- Respect robots.txt and site terms.
- Use conditional requests and caching to avoid re-fetching.
- Keep a more realistic session/cookie flow if allowed.
- Handle 429/403 gracefully and back off.
Important note
If the site is intentionally blocking automation, the right approach is usually to obtain permission, use an official API, or limit crawling to what’s permitted. I can help you diagnose the likely cause if you share:
- the HTTP status codes you’re seeing,
- the request rate/concurrency,
- and a sample of the response headers/body from a blocked request.
Brands