Prompt

How do I build an enterprise scraper with retries, proxies, and parsing?

Technology / Data infrastructure1 observationLast seen Sep 24, 2026

Latest observation

Sep 24, 2026OpenAI APIWeb search: off

To build an enterprise-grade scraper with retries, proxies, and parsing, think in terms of a pipeline rather than a single script:

  1. Fetch pages reliably
  2. Retry failures intelligently
  3. Rotate proxies and manage IP reputation
  4. Parse HTML into structured data
  5. Persist results with observability and deduplication
  6. Operate it safely at scale

Here’s a practical blueprint.


1) Core architecture

A robust scraper usually has these components:

  • Scheduler / queue: decides what URLs to fetch
  • Fetcher: downloads pages
  • Retry policy: handles transient failures
  • Proxy manager: selects and validates proxies
  • Parser: extracts structured data from HTML/JSON
  • Storage: database, blob storage, or queue
  • Monitoring: logs, metrics, alerts
  • Rate limiting / throttling: avoids overload and bans

A common flow:

URL queue -> fetcher -> retry/proxy middleware -> parser -> validation -> storage


2) Build the fetcher with retries

Use retries for:

  • timeouts
  • 429 rate limits
  • 5xx server errors
  • temporary DNS/network failures

Use exponential backoff with jitter to avoid retry storms.

Example (Python + requests)

import random
import time
import requests
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry

def make_session():
    session = requests.Session()
    retry = Retry(
        total=5,
        backoff_factor=1.5,
        status_forcelist=[429, 500, 502, 503, 504],
        allowed_methods=["GET", "HEAD"],
        raise_on_status=False,
    )
    adapter = HTTPAdapter(max_retries=retry)
    session.mount("http://", adapter)
    session.mount("https://", adapter)
    return session

def fetch(url, session=None, headers=None, timeout=20, proxies=None):
    session = session or make_session()
    headers = headers or {
        "User-Agent": "Mozilla/5.0 (compatible; EnterpriseScraper/1.0)"
    }
    resp = session.get(url, headers=headers, timeout=timeout, proxies=proxies)
    resp.raise_for_status()
    return resp.text

Notes

  • Avoid infinite retries.
  • Treat 429 specially with longer backoff.
  • Use different retry rules for GET vs POST.

3) Add proxy support

For enterprise scraping, proxies help with:

  • IP rotation
  • geo targeting
  • resilience
  • traffic distribution

Proxy strategy

Use a proxy pool with:

  • health checks
  • latency tracking
  • failure scoring
  • ban detection
  • sticky sessions when needed

Simple proxy usage

proxies = {
    "http": "http://user:pass@proxy-host:8080",
    "https": "http://user:pass@proxy-host:8080",
}

html = fetch("https://example.com", proxies=proxies)

Better proxy manager idea

Maintain a list of proxies and choose one based on:

  • success rate
  • current error rate
  • response time
  • target domain affinity

Pseudo-logic:

class ProxyManager:
    def __init__(self, proxies):
        self.proxies = proxies
        self.stats = {p: {"ok": 0, "fail": 0, "latency": []} for p in proxies}

    def pick(self):
        # pick healthiest proxy
        return min(self.proxies, key=lambda p: self.score(p))

    def score(self, p):
        s = self.stats[p]
        fail_rate = (s["fail"] + 1) / (s["ok"] + s["fail"] + 2)
        avg_latency = sum(s["latency"]) / len(s["latency"]) if s["latency"] else 1
        return fail_rate * 10 + avg_latency

    def mark_ok(self, p, latency):
        self.stats[p]["ok"] += 1
        self.stats[p]["latency"].append(latency)

    def mark_fail(self, p):
        self.stats[p]["fail"] += 1

4) Parse reliably

Parsing is often the hardest part because HTML changes.

Use:

  • BeautifulSoup / lxml for HTML
  • parsel / cssselect / XPath for more precise extraction
  • json.loads if the site embeds JSON
  • schema validation to ensure extracted data is sane

Example using BeautifulSoup

from bs4 import BeautifulSoup

def parse_product(html):
    soup = BeautifulSoup(html, "html.parser")

    title = soup.select_one("h1.product-title")
    price = soup.select_one(".price")
    sku = soup.select_one("[data-sku]")

    return {
        "title": title.get_text(strip=True) if title else None,
        "price": price.get_text(strip=True) if price else None,
        "sku": sku["data-sku"] if sku and sku.has_attr("data-sku") else None,
    }

Best practices for parsing

  • Prefer stable attributes: data-*, IDs, semantic classes
  • Build parser tests with saved HTML fixtures
  • Handle missing fields gracefully
  • Normalize values early:
    • strip currency symbols
    • convert dates to ISO-8601
    • parse numbers safely

5) Make it resilient

Retry only what’s safe

  • Safe to retry: GET, idempotent reads
  • Be careful with:
    • form submissions
    • actions that mutate data

Add deduplication

Store:

  • URL hashes
  • content hashes
  • canonical IDs from the site

Handle bans and blocks

Detect:

  • captchas
  • unusual redirects
  • repeated 403/429
  • honeypot pages

When detected:

  • slow down
  • rotate proxy
  • change fingerprint if appropriate
  • mark target as blocked

Timeouts

Set explicit timeouts for:

  • connect
  • read
  • total request

6) Use a queue and workers for scale

For large-scale scraping:

  • Put URLs into a queue
  • Run multiple workers
  • Limit concurrency per domain
  • Use per-domain rate limits

Good tools:

  • Celery + Redis/RabbitMQ
  • RQ
  • Kafka
  • SQS
  • asyncio + aiohttp for high concurrency

Domain-aware throttling

Don’t hammer one host. Implement:

  • N requests per second per domain
  • randomized delays
  • backpressure when error rate rises

7) Monitoring and observability

Track:

  • request success/failure rates
  • status code distribution
  • average latency
  • proxy health
  • parse success rate
  • record counts
  • block/captcha rate

Log structured events:

{
  "url": "https://example.com/item/123",
  "status": 200,
  "proxy": "proxy-7",
  "latency_ms": 842,
  "parsed": true,
  "error": null
}

Alerts:

  • spike in 403/429
  • parse failures after site changes
  • proxy failure surge
  • queue backlog growth

8) A simple end-to-end example

import time
import requests
from bs4 import BeautifulSoup

class Scraper:
    def __init__(self, proxies=None):
        self.session = requests.Session()
        self.proxies = proxies or []

    def get_proxy(self):
        return self.proxies[0] if self.proxies else None

    def fetch(self, url):
        proxy = self.get_proxy()
        proxies = {"http": proxy, "https": proxy} if proxy else None

        headers = {"User-Agent": "Mozilla/5.0 (compatible; EnterpriseScraper/1.0)"}
        for attempt in range(5):
            try:
                r = self.session.get(url, headers=headers, proxies=proxies, timeout=20)
                if r.status_code in (429, 500, 502, 503, 504):
                    time.sleep((2 ** attempt) + 0.1)
                    continue
                r.raise_for_status()
                return r.text
            except requests.RequestException:
                if attempt == 4:
                    raise
                time.sleep((2 ** attempt) + 0.1)

    def parse(self, html):
        soup = BeautifulSoup(html, "html.parser")
        return {
            "title": soup.select_one("title").get_text(strip=True) if soup.select_one("title") else None,
        }

    def run(self, url):
        html = self.fetch(url)
        return self.parse(html)

scraper = Scraper()
print(scraper.run("https://example.com"))

9) Recommended stack

Small-to-medium

  • Python
  • requests or httpx
  • BeautifulSoup + lxml
  • SQLite/Postgres
  • Redis for queues

Large-scale

  • aiohttp or httpx async
  • Celery / Kafka / SQS
  • Postgres + object storage
  • Prometheus + Grafana
  • Proxy pool with health scoring

10) Production checklist

Before shipping, make sure you have:

  • Retries with backoff + jitter
  • Proxy rotation and health checks
  • Per-domain rate limiting
  • Timeouts everywhere
  • Parser tests with fixtures
  • Structured logging
  • Metrics and alerts
  • Deduplication
  • Storage schema
  • Error handling for blocks and captcha pages
  • Compliance review for target sites and legal constraints

11) Key design advice

  • Don’t make parsing depend on brittle CSS classes only.
  • Separate fetch, parse, and storage so each can evolve independently.
  • Keep proxy logic isolated from business logic.
  • Measure everything; scraper reliability is mostly an operations problem.

If you want, I can give you:

  1. a production-ready Python template,
  2. an async scraper with httpx/aiohttp, or
  3. a Scrapy-based enterprise architecture.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.