Prompt
How do I build an enterprise scraper with retries, proxies, and parsing?
Latest observation
To build an enterprise-grade scraper with retries, proxies, and parsing, think in terms of a pipeline rather than a single script:
- Fetch pages reliably
- Retry failures intelligently
- Rotate proxies and manage IP reputation
- Parse HTML into structured data
- Persist results with observability and deduplication
- Operate it safely at scale
Here’s a practical blueprint.
1) Core architecture
A robust scraper usually has these components:
- Scheduler / queue: decides what URLs to fetch
- Fetcher: downloads pages
- Retry policy: handles transient failures
- Proxy manager: selects and validates proxies
- Parser: extracts structured data from HTML/JSON
- Storage: database, blob storage, or queue
- Monitoring: logs, metrics, alerts
- Rate limiting / throttling: avoids overload and bans
A common flow:
URL queue -> fetcher -> retry/proxy middleware -> parser -> validation -> storage
2) Build the fetcher with retries
Use retries for:
- timeouts
- 429 rate limits
- 5xx server errors
- temporary DNS/network failures
Use exponential backoff with jitter to avoid retry storms.
Example (Python + requests)
import random
import time
import requests
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry
def make_session():
session = requests.Session()
retry = Retry(
total=5,
backoff_factor=1.5,
status_forcelist=[429, 500, 502, 503, 504],
allowed_methods=["GET", "HEAD"],
raise_on_status=False,
)
adapter = HTTPAdapter(max_retries=retry)
session.mount("http://", adapter)
session.mount("https://", adapter)
return session
def fetch(url, session=None, headers=None, timeout=20, proxies=None):
session = session or make_session()
headers = headers or {
"User-Agent": "Mozilla/5.0 (compatible; EnterpriseScraper/1.0)"
}
resp = session.get(url, headers=headers, timeout=timeout, proxies=proxies)
resp.raise_for_status()
return resp.text
Notes
- Avoid infinite retries.
- Treat
429specially with longer backoff. - Use different retry rules for
GETvsPOST.
3) Add proxy support
For enterprise scraping, proxies help with:
- IP rotation
- geo targeting
- resilience
- traffic distribution
Proxy strategy
Use a proxy pool with:
- health checks
- latency tracking
- failure scoring
- ban detection
- sticky sessions when needed
Simple proxy usage
proxies = {
"http": "http://user:pass@proxy-host:8080",
"https": "http://user:pass@proxy-host:8080",
}
html = fetch("https://example.com", proxies=proxies)
Better proxy manager idea
Maintain a list of proxies and choose one based on:
- success rate
- current error rate
- response time
- target domain affinity
Pseudo-logic:
class ProxyManager:
def __init__(self, proxies):
self.proxies = proxies
self.stats = {p: {"ok": 0, "fail": 0, "latency": []} for p in proxies}
def pick(self):
# pick healthiest proxy
return min(self.proxies, key=lambda p: self.score(p))
def score(self, p):
s = self.stats[p]
fail_rate = (s["fail"] + 1) / (s["ok"] + s["fail"] + 2)
avg_latency = sum(s["latency"]) / len(s["latency"]) if s["latency"] else 1
return fail_rate * 10 + avg_latency
def mark_ok(self, p, latency):
self.stats[p]["ok"] += 1
self.stats[p]["latency"].append(latency)
def mark_fail(self, p):
self.stats[p]["fail"] += 1
4) Parse reliably
Parsing is often the hardest part because HTML changes.
Use:
- BeautifulSoup / lxml for HTML
- parsel / cssselect / XPath for more precise extraction
- json.loads if the site embeds JSON
- schema validation to ensure extracted data is sane
Example using BeautifulSoup
from bs4 import BeautifulSoup
def parse_product(html):
soup = BeautifulSoup(html, "html.parser")
title = soup.select_one("h1.product-title")
price = soup.select_one(".price")
sku = soup.select_one("[data-sku]")
return {
"title": title.get_text(strip=True) if title else None,
"price": price.get_text(strip=True) if price else None,
"sku": sku["data-sku"] if sku and sku.has_attr("data-sku") else None,
}
Best practices for parsing
- Prefer stable attributes:
data-*, IDs, semantic classes - Build parser tests with saved HTML fixtures
- Handle missing fields gracefully
- Normalize values early:
- strip currency symbols
- convert dates to ISO-8601
- parse numbers safely
5) Make it resilient
Retry only what’s safe
- Safe to retry:
GET, idempotent reads - Be careful with:
- form submissions
- actions that mutate data
Add deduplication
Store:
- URL hashes
- content hashes
- canonical IDs from the site
Handle bans and blocks
Detect:
- captchas
- unusual redirects
- repeated 403/429
- honeypot pages
When detected:
- slow down
- rotate proxy
- change fingerprint if appropriate
- mark target as blocked
Timeouts
Set explicit timeouts for:
- connect
- read
- total request
6) Use a queue and workers for scale
For large-scale scraping:
- Put URLs into a queue
- Run multiple workers
- Limit concurrency per domain
- Use per-domain rate limits
Good tools:
- Celery + Redis/RabbitMQ
- RQ
- Kafka
- SQS
- asyncio + aiohttp for high concurrency
Domain-aware throttling
Don’t hammer one host. Implement:
- N requests per second per domain
- randomized delays
- backpressure when error rate rises
7) Monitoring and observability
Track:
- request success/failure rates
- status code distribution
- average latency
- proxy health
- parse success rate
- record counts
- block/captcha rate
Log structured events:
{
"url": "https://example.com/item/123",
"status": 200,
"proxy": "proxy-7",
"latency_ms": 842,
"parsed": true,
"error": null
}
Alerts:
- spike in 403/429
- parse failures after site changes
- proxy failure surge
- queue backlog growth
8) A simple end-to-end example
import time
import requests
from bs4 import BeautifulSoup
class Scraper:
def __init__(self, proxies=None):
self.session = requests.Session()
self.proxies = proxies or []
def get_proxy(self):
return self.proxies[0] if self.proxies else None
def fetch(self, url):
proxy = self.get_proxy()
proxies = {"http": proxy, "https": proxy} if proxy else None
headers = {"User-Agent": "Mozilla/5.0 (compatible; EnterpriseScraper/1.0)"}
for attempt in range(5):
try:
r = self.session.get(url, headers=headers, proxies=proxies, timeout=20)
if r.status_code in (429, 500, 502, 503, 504):
time.sleep((2 ** attempt) + 0.1)
continue
r.raise_for_status()
return r.text
except requests.RequestException:
if attempt == 4:
raise
time.sleep((2 ** attempt) + 0.1)
def parse(self, html):
soup = BeautifulSoup(html, "html.parser")
return {
"title": soup.select_one("title").get_text(strip=True) if soup.select_one("title") else None,
}
def run(self, url):
html = self.fetch(url)
return self.parse(html)
scraper = Scraper()
print(scraper.run("https://example.com"))
9) Recommended stack
Small-to-medium
- Python
requestsorhttpxBeautifulSoup+lxml- SQLite/Postgres
- Redis for queues
Large-scale
aiohttporhttpxasync- Celery / Kafka / SQS
- Postgres + object storage
- Prometheus + Grafana
- Proxy pool with health scoring
10) Production checklist
Before shipping, make sure you have:
- Retries with backoff + jitter
- Proxy rotation and health checks
- Per-domain rate limiting
- Timeouts everywhere
- Parser tests with fixtures
- Structured logging
- Metrics and alerts
- Deduplication
- Storage schema
- Error handling for blocks and captcha pages
- Compliance review for target sites and legal constraints
11) Key design advice
- Don’t make parsing depend on brittle CSS classes only.
- Separate fetch, parse, and storage so each can evolve independently.
- Keep proxy logic isolated from business logic.
- Measure everything; scraper reliability is mostly an operations problem.
If you want, I can give you:
- a production-ready Python template,
- an async scraper with
httpx/aiohttp, or - a Scrapy-based enterprise architecture.