Prompt
How do I crawl pages with rate limits and retries?
Latest observation
Oct 4, 2026
OpenAI APIWeb search: off
To crawl pages safely with rate limits and retries, you want to do three things:
- Limit request pace so you don’t overwhelm the server.
- Retry transient failures like 429/500/502/503/504.
- Back off intelligently when the site asks you to slow down.
Core approach
1) Add a rate limiter
Use a fixed delay or token bucket so requests stay under a target rate.
- Simple:
sleep(1 / requests_per_second)between requests - Better: token bucket or leaky bucket
- Respect
Retry-Afterheaders when present
2) Retry only transient errors
Retry on:
429 Too Many Requests500, 502, 503, 504- network timeouts / connection resets
Do not retry immediately on permanent errors like 404 or 403 unless you have a special reason.
3) Use exponential backoff with jitter
Example delays:
- attempt 1: 1s
- attempt 2: 2s
- attempt 3: 4s
- attempt 4: 8s
Add random jitter so many clients don’t retry at the same moment.
4) Cap retries
Usually 3–5 retries is enough. After that, log and move on.
5) Crawl politely
- Identify your user agent
- Check
robots.txtif appropriate - Reduce concurrency
- Cache responses if possible
Example in Python (requests + urllib3 retry)
import time
import random
import requests
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry
session = requests.Session()
retry = Retry(
total=5,
backoff_factor=1, # exponential backoff
status_forcelist=[429, 500, 502, 503, 504],
allowed_methods=["GET", "HEAD"],
respect_retry_after_header=True,
)
adapter = HTTPAdapter(max_retries=retry)
session.mount("http://", adapter)
session.mount("https://", adapter)
session.headers.update({
"User-Agent": "MyCrawler/1.0 (+contact@example.com)"
})
urls = ["https://example.com/page1", "https://example.com/page2"]
for url in urls:
try:
response = session.get(url, timeout=15)
response.raise_for_status()
print(url, len(response.text))
time.sleep(0.5) # simple rate limit
except requests.RequestException as e:
print(f"Failed: {url} -> {e}")
Example with manual retries + backoff
import time
import random
import requests
def fetch(url, session, max_retries=5):
for attempt in range(max_retries):
try:
r = session.get(url, timeout=15)
if r.status_code == 200:
return r.text
if r.status_code in (429, 500, 502, 503, 504):
retry_after = r.headers.get("Retry-After")
if retry_after:
sleep_for = int(retry_after)
else:
sleep_for = (2 ** attempt) + random.uniform(0, 1)
time.sleep(sleep_for)
continue
r.raise_for_status()
except requests.RequestException:
sleep_for = (2 ** attempt) + random.uniform(0, 1)
time.sleep(sleep_for)
raise RuntimeError(f"Failed after {max_retries} retries: {url}")
If you’re crawling many pages
Use concurrency carefully:
- Start with 1–3 concurrent workers
- Add per-host rate limiting
- Don’t retry all failed URLs at once
- Queue retries with delay
Good defaults
- Requests per second: 1–2 per host
- Retries: 3–5
- Timeout: 10–30 seconds
- Backoff: exponential with jitter
- Concurrency: low unless the site explicitly allows more
If you want, I can also show:
- a Scrapy configuration for rate limits and retries
- an asyncio/httpx version
- a per-domain crawler architecture with a queue and retry scheduler