Prompt
What should I use to crawl a million pages a month?
Latest observation
For 1 million pages/month, the best choice depends on how heavy the pages are, how fresh the data must be, and whether you need rendering/login/JS. Roughly, that’s only about 23 pages/minute on average, so the scale itself isn’t huge—but reliability matters more than raw speed.
Good options
1. Simple static pages: Python + Scrapy
Best if pages are mostly HTML and don’t require JavaScript rendering.
- Very efficient and lightweight
- Built-in throttling, retries, deduping
- Easy to run on a single VM or a small cluster
- Good ecosystem for parsing/storing data
Use this if: you want control, low cost, and pages are crawl-friendly.
2. JavaScript-heavy sites: Playwright (or Puppeteer) + queue
Best if pages need browser rendering.
- Handles modern sites well
- More resource-intensive than Scrapy
- Usually pair with a task queue like Celery, RabbitMQ, SQS, or Kafka
Use this if: you need to render pages, interact with buttons, or handle anti-bot-ish frontends.
3. Managed crawling platform: Apify, Zyte, Bright Data, Oxylabs
Best if you want less ops work.
- Proxy management included
- Scaling and retries are easier
- Often useful for hard targets
- Higher cost, less control
Use this if: you care more about reliability and speed to production than running infrastructure yourself.
My practical recommendation
If you’re starting from scratch:
- Static pages:
Scrapy + PostgreSQL/S3 + Redis queue - Dynamic pages:
Playwright + queue + rotating proxies - At scale / hard sites: managed service like Zyte or Bright Data
Infrastructure estimate
For 1M pages/month, you can often do this with:
- 1–3 modest servers for static crawling
- More if using browsers (Playwright can be 5–20x heavier)
- Add:
- Redis for scheduling/dupe filtering
- PostgreSQL for metadata
- S3/GCS for raw HTML if needed
Important caveats
Make sure you consider:
- robots.txt / site terms
- rate limiting and politeness
- retries and backoff
- deduplication
- monitoring for blocks/errors
If you want, I can recommend a setup based on:
- static vs JS pages,
- how many domains, and
- whether you need full HTML or just extracted fields.