Prompt
I'm building a bulk scraper for competitor sites with retries and deduping, what stack makes sense?
Latest observation
Oct 4, 2026
OpenAI APIWeb search: off
A solid bulk-scraping stack depends on scale, target sites, and how much anti-bot friction you expect, but for retryable, deduped, large-volume scraping I’d usually recommend this shape:
Good default stack
1) Orchestration / jobs
- Python for scraping logic
- Celery + Redis or RQ for async job queues
- If you need higher reliability / more workflow control: Temporal or Airflow
2) HTTP fetching
- httpx or aiohttp for fast concurrent requests
- Use asyncio for concurrency
- Add a retry library like tenacity
3) Browser automation for JS-heavy sites
- Playwright is usually the best default
- Use it only when needed; keep most pages on plain HTTP to stay fast and cheaper
4) Deduping / idempotency
- Postgres as the source of truth
- Unique constraints on:
- normalized URL
- site + product ID
- content hash
- For high-scale approximate dedupe, add Redis Bloom filter or a Bloom filter in-memory layer
5) Storage
- Postgres for metadata and state
- S3/GCS for raw HTML, screenshots, and extracted artifacts
- Optionally Elasticsearch/OpenSearch if you need search over captured pages
6) Parsing / extraction
- BeautifulSoup + lxml for static HTML
- selectolax if you want faster parsing
- XPath/CSS selectors with site-specific extractors
7) Observability
- Prometheus + Grafana for metrics
- Sentry for exception tracking
- Structured logs with JSON logging
Recommended architecture
Ingestion flow
- URL discovered or imported
- Normalize and dedupe
- Queue fetch job
- Fetch with retries/backoff
- If blocked or JS required, route to Playwright
- Parse content
- Persist raw + extracted data
- Mark status and store hash for dedupe
Retry strategy
- Retry on:
- 429
- 5xx
- connection timeouts
- transient DNS/socket failures
- Exponential backoff with jitter
- Cap retries per URL to avoid infinite loops
- Keep a dead-letter queue for failed jobs
Dedupe strategy
Use multiple layers:
- URL normalization: remove tracking params, canonicalize scheme/host/path
- Content hashing: hash meaningful body content after boilerplate removal
- Entity dedupe: product/company IDs if available
- Seen-URL store: Postgres unique index or Redis set for fast checks
If you expect anti-bot measures
Then add:
- Proxy rotation through a reputable provider
- Per-domain rate limiting
- Cookie/session management
- User-agent rotation, but don’t rely on that alone
- Separate queues per domain to avoid hammering a single site
Best “practical” stack by scale
Small/medium scale
- Python
- httpx + asyncio
- tenacity
- BeautifulSoup/lxml
- Postgres
- Redis
- Celery
- Playwright for fallback
Larger scale
- Python workers
- Kafka or RabbitMQ instead of plain Redis queues
- Temporal for orchestration
- Postgres + S3
- Playwright isolated in a separate worker pool
- Redis Bloom / approximate dedupe
- OpenSearch for retrieval
My opinionated recommendation
If you want something that’s fast to build and production-friendly, start with:
- Python
- httpx + asyncio
- Playwright
- Postgres
- Redis
- Celery
- tenacity
- S3
- Prometheus/Grafana
That gives you a very good balance of speed, maintainability, and scale.
One important note
Make sure you’re scraping in a way that respects the target sites’ terms, robots rules where applicable, and rate limits. For competitor monitoring, it’s also worth checking legal/compliance boundaries in your jurisdiction.
If you want, I can also give you:
- a reference architecture diagram,
- a minimal production-ready folder structure, or
- a sample Python crawler with retries + dedupe.