Prompt

How do I crawl pages with rate limits and retries?

Technology / Cloud Infrastructure1 observationLast seen Oct 4, 2026

Latest observation

Oct 4, 2026OpenAI APIWeb search: off

To crawl pages safely with rate limits and retries, you want to do three things:

  1. Limit request pace so you don’t overwhelm the server.
  2. Retry transient failures like 429/500/502/503/504.
  3. Back off intelligently when the site asks you to slow down.

Core approach

1) Add a rate limiter

Use a fixed delay or token bucket so requests stay under a target rate.

  • Simple: sleep(1 / requests_per_second) between requests
  • Better: token bucket or leaky bucket
  • Respect Retry-After headers when present

2) Retry only transient errors

Retry on:

  • 429 Too Many Requests
  • 500, 502, 503, 504
  • network timeouts / connection resets

Do not retry immediately on permanent errors like 404 or 403 unless you have a special reason.

3) Use exponential backoff with jitter

Example delays:

  • attempt 1: 1s
  • attempt 2: 2s
  • attempt 3: 4s
  • attempt 4: 8s

Add random jitter so many clients don’t retry at the same moment.

4) Cap retries

Usually 3–5 retries is enough. After that, log and move on.

5) Crawl politely

  • Identify your user agent
  • Check robots.txt if appropriate
  • Reduce concurrency
  • Cache responses if possible

Example in Python (requests + urllib3 retry)

import time
import random
import requests
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry

session = requests.Session()

retry = Retry(
    total=5,
    backoff_factor=1,  # exponential backoff
    status_forcelist=[429, 500, 502, 503, 504],
    allowed_methods=["GET", "HEAD"],
    respect_retry_after_header=True,
)

adapter = HTTPAdapter(max_retries=retry)
session.mount("http://", adapter)
session.mount("https://", adapter)

session.headers.update({
    "User-Agent": "MyCrawler/1.0 (+contact@example.com)"
})

urls = ["https://example.com/page1", "https://example.com/page2"]

for url in urls:
    try:
        response = session.get(url, timeout=15)
        response.raise_for_status()
        print(url, len(response.text))
        time.sleep(0.5)  # simple rate limit
    except requests.RequestException as e:
        print(f"Failed: {url} -> {e}")

Example with manual retries + backoff

import time
import random
import requests

def fetch(url, session, max_retries=5):
    for attempt in range(max_retries):
        try:
            r = session.get(url, timeout=15)
            if r.status_code == 200:
                return r.text
            if r.status_code in (429, 500, 502, 503, 504):
                retry_after = r.headers.get("Retry-After")
                if retry_after:
                    sleep_for = int(retry_after)
                else:
                    sleep_for = (2 ** attempt) + random.uniform(0, 1)
                time.sleep(sleep_for)
                continue
            r.raise_for_status()
        except requests.RequestException:
            sleep_for = (2 ** attempt) + random.uniform(0, 1)
            time.sleep(sleep_for)

    raise RuntimeError(f"Failed after {max_retries} retries: {url}")

If you’re crawling many pages

Use concurrency carefully:

  • Start with 1–3 concurrent workers
  • Add per-host rate limiting
  • Don’t retry all failed URLs at once
  • Queue retries with delay

Good defaults

  • Requests per second: 1–2 per host
  • Retries: 3–5
  • Timeout: 10–30 seconds
  • Backoff: exponential with jitter
  • Concurrency: low unless the site explicitly allows more

If you want, I can also show:

  • a Scrapy configuration for rate limits and retries
  • an asyncio/httpx version
  • a per-domain crawler architecture with a queue and retry scheduler

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.