Prompt
How do I build a proxy-backed scraper for recurring data pulls?
Latest observation
To build a proxy-backed scraper for recurring data pulls, think of it as a small data pipeline with four parts:
- Target fetcher
- Proxy layer
- Parsing / normalization
- Scheduler + storage + monitoring
Here’s a practical blueprint.
1) Start with the right architecture
A robust recurring scraper usually looks like this:
- Scheduler triggers jobs on a cadence
- Worker fetches pages through proxies
- Parser extracts the fields you need
- Database / warehouse stores snapshots or upserts
- Logger/monitor tracks failures, bans, and latency
Typical flow:
Cron / Airflow / Celery / queue
↓
Scraper worker
↓
Proxy rotation
↓
Target website
↓
HTML/JSON parse
↓
Clean + normalize
↓
Store results
2) Choose the right proxy strategy
Proxy types
- Datacenter proxies: cheaper, faster, more likely to be blocked
- Residential proxies: harder to block, more expensive
- Mobile proxies: best evasion, highest cost
Rotation strategies
Use one of these:
- Round-robin: simple rotation through a list
- Sticky sessions: keep the same IP per target/session for a while
- Adaptive rotation: change proxy when you see bans, 403s, CAPTCHA, or rate-limit responses
For recurring pulls, sticky sessions per domain or account often work best unless the site is very aggressive.
3) Make the scraper resilient
Handle failures gracefully
Your worker should:
- Retry transient errors
- Detect block pages
- Back off on rate limits
- Switch proxies when needed
- Save partial progress so a run can resume
Use headers and pacing
Even with proxies, websites can flag:
- identical request bursts
- missing browser-like headers
- no cookies/session continuity
Use:
- realistic
User-Agent Accept-LanguageAccept,Refererif applicable- modest concurrency
- randomized delays
4) Basic Python example with rotating proxies
Here’s a simple requests-based pattern:
import random
import time
import requests
PROXIES = [
"http://user:pass@proxy1.example.com:8000",
"http://user:pass@proxy2.example.com:8000",
"http://user:pass@proxy3.example.com:8000",
]
HEADERS = {
"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 "
"(KHTML, like Gecko) Chrome/124.0.0.0 Safari/537.36",
"Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8",
"Accept-Language": "en-US,en;q=0.9",
}
def fetch(url, max_retries=5):
last_err = None
for attempt in range(max_retries):
proxy = {"http": random.choice(PROXIES), "https": random.choice(PROXIES)}
try:
resp = requests.get(url, headers=HEADERS, proxies=proxy, timeout=20)
if resp.status_code in (403, 429):
raise RuntimeError(f"Blocked/rate-limited: {resp.status_code}")
resp.raise_for_status()
return resp.text
except Exception as e:
last_err = e
time.sleep(2 ** attempt) # exponential backoff
raise last_err
html = fetch("https://example.com/data")
print(html[:500])
Notes
- For real jobs, don’t pick
httpandhttpsproxies independently at random if they’re meant to represent the same endpoint. - Add logging for each attempt, status code, proxy used, and duration.
5) If the site is JavaScript-heavy, use a browser automation tool
If the content is rendered client-side, use:
- Playwright (recommended)
- Selenium
- Puppeteer (Node.js)
With Playwright, you can route traffic via proxy and reuse sessions.
Example:
from playwright.sync_api import sync_playwright
proxy = {
"server": "http://proxy1.example.com:8000",
"username": "user",
"password": "pass"
}
with sync_playwright() as p:
browser = p.chromium.launch(headless=True, proxy=proxy)
page = browser.new_page()
page.goto("https://example.com/data", wait_until="networkidle")
content = page.content()
print(content[:500])
browser.close()
Use browser automation when:
- data is loaded after page render
- API calls are signed or hidden behind JS
- anti-bot checks depend on browser behavior
6) Store data in a way that supports recurring pulls
For recurring scraping, storage design matters.
Common patterns
- Append-only snapshots: useful for tracking changes over time
- Upserts by unique key: best when records have stable IDs
- Versioned history: track every change with timestamps
Suggested fields
Store:
- source URL
- fetch timestamp
- proxy region/provider
- HTTP status
- parsed record ID
- raw payload or hash
- normalized fields
- run ID
Example table columns:
record_id | name | price | status | source_url | fetched_at | run_id
7) Build a scheduler for recurring runs
Use one of:
- cron for simple schedules
- Celery beat
- Airflow
- Prefect
- Dagster
For example, cron:
0 */6 * * * /usr/bin/python3 /app/scrape_job.py >> /var/log/scraper.log 2>&1
For many targets, a workflow tool is better because you can:
- fan out jobs
- retry selectively
- track dependencies
- store run history
8) Add observability from day one
Track:
- success rate
- response times
- block/rate-limit rates
- proxy health
- parse success rate
- number of records extracted
Useful logs:
- target URL
- proxy ID
- status code
- elapsed time
- error type
- records extracted
If a proxy starts failing often, remove it automatically.
9) Use a queue if volume grows
If you’re pulling data from many pages regularly, use:
- Redis queue
- RabbitMQ
- SQS
- Kafka if very large scale
This gives you:
- distributed workers
- retries
- deduplication
- better throughput control
10) Best practices for reliability
- Respect site terms and robots rules where applicable
- Keep concurrency moderate
- Cache when possible
- Deduplicate URLs and records
- Use idempotent writes
- Save raw responses for debugging
- Rotate proxies only when needed, not on every request if sticky sessions help
- Avoid hammering the same endpoint at fixed intervals
11) A good production pattern
For recurring pulls, a strong baseline is:
- Playwright or requests
- Proxy pool with health scoring
- Queue-based job distribution
- Retry/backoff
- Structured logs
- Postgres for storage
- Cron/Airflow for scheduling
12) Example production flow
- Scheduler creates jobs for target pages
- Worker claims a job
- Worker selects a healthy proxy
- Fetch page
- Detect block/rate limit
- Retry with a different proxy if needed
- Parse and normalize fields
- Upsert into DB
- Emit metrics and logs
- Mark job complete
If you want, I can also give you:
- a complete Python scraper template
- a Playwright + rotating proxy implementation
- a Postgres schema for recurring pulls
- or a Dockerized production setup