Prompt
I'm building a crawler for thousands of domains, what should I use?
Latest observation
Oct 4, 2026
OpenAI APIWeb search: off
For thousands of domains, you want something that’s built for high-throughput, concurrent, fault-tolerant crawling, not a simple single-threaded scraper.
Best choices depending on your needs
1) Scrapy (Python) — best general-purpose option
Use this if you want:
- Fast crawling at scale
- Built-in concurrency
- Automatic retries, throttling, link following, deduping
- Good support for custom pipelines and exporters
Why it’s good:
- Mature and widely used
- Easy to distribute later if needed
- Great for structured extraction
Typical stack:
- Scrapy for crawling
- Redis or Kafka for distributed scheduling
- PostgreSQL / S3 / Elasticsearch for storage
2) Playwright or Selenium — only if you need heavy JavaScript rendering
Use this if:
- Pages are JS-rendered
- Content is hidden behind client-side app logic
- You need to interact with the site
Downside:
- Much slower and more expensive than plain HTTP crawling
- Not ideal for “thousands of domains” unless only a small subset needs rendering
Best practice:
- Use Scrapy first
- Fall back to Playwright only for sites that require it
3) StormCrawler / Apache Nutch — for very large distributed crawling
Use this if:
- You’re crawling at internet scale
- You need a distributed architecture across many machines
- You already use Java / Hadoop / Kafka ecosystems
Downside:
- More complex to operate
- Overkill unless you’re truly doing large-scale crawling
4) Custom crawler with asyncio + aiohttp/httpx — if you need full control
Use this if:
- You want a lightweight custom solution
- Your crawling logic is very specific
- You’re comfortable building dedupe, retries, politeness, robots handling, and queues yourself
This can be very fast, but you’ll need to implement:
- Rate limiting per domain
- Retry/backoff
- URL normalization
- Deduplication
- Persistence
- Monitoring
My recommendation
If you’re just getting started:
Use Scrapy.
It gives you the best balance of:
- speed
- maturity
- ease of development
- scalability
If you need JS rendering:
Scrapy + Playwright
If you need distributed crawling across many machines:
Scrapy + Redis/Kafka, or consider StormCrawler if your team is Java-heavy.
Important features your crawler should have
For thousands of domains, make sure it supports:
- Per-domain concurrency limits
- Politeness delays
- Robots.txt handling
- Retries with exponential backoff
- URL canonicalization and deduplication
- Timeouts and error handling
- Proxy support if needed
- Logging/metrics for failed domains and crawl depth
Quick rule of thumb
- Mostly static HTML → Scrapy
- JS-heavy sites → Scrapy + Playwright
- Massively distributed / enterprise scale → StormCrawler/Nutch
- Need custom low-level control → asyncio/httpx
If you want, I can suggest a specific architecture for crawling thousands of domains efficiently, including queues, storage, retries, and how to avoid getting blocked.