Prompt
How do I crawl millions of pages without my own infra?
Latest observation
If you want to crawl millions of pages without running your own infrastructure, the practical answer is: don’t try to “scrape the web” directly from one machine. Use a combination of managed crawling tools, public datasets, and serverless/batch compute.
Best options
1) Use a managed crawling platform
These handle queueing, retries, scaling, proxies, and sometimes rendering.
Examples:
- Apify
- Bright Data Web Scraper / Data Collector
- Zyte
- ScrapingBee
- Oxylabs
- Common Crawl if you just need already-crawled web data
Good for:
- large-scale scraping jobs
- JS-heavy sites
- proxy management
- avoiding ops work
Tradeoff:
- can get expensive at very large scale
- still need to respect site policies and legal constraints
2) Use serverless / batch compute
If you need custom crawling logic but not your own servers, run the crawler in:
- AWS Lambda
- AWS ECS/Fargate
- Google Cloud Run
- Google Batch
- Azure Container Apps
- Fly.io Machines
- Render Background Workers
- Temporal Cloud for orchestration
Typical pattern:
- a queue of URLs
- workers pull URLs in parallel
- store results in object storage / database
- throttle per domain
Good for:
- custom logic
- pay-per-use
- no server maintenance
Tradeoff:
- you still design the crawler architecture
- need careful rate limiting and deduplication
3) Use existing web datasets
If your goal is data extraction rather than live crawling, this is often the cheapest route.
Sources:
- Common Crawl
- Internet Archive
- Wikidata / Wikipedia dumps
- domain-specific datasets
- commercial data providers
Good for:
- huge scale
- historical snapshots
- avoiding bot defenses entirely
Tradeoff:
- data may be stale or incomplete
- not all sites are represented
A scalable architecture without owning infra
A common approach:
- Seed URLs
- Put them in a managed queue
- Use serverless workers to fetch pages
- Store HTML/metadata in S3/GCS/Azure Blob
- Extract links and enqueue new URLs
- Deduplicate with a URL hash set
- Enforce per-domain politeness
- Export cleaned data to your warehouse or search index
Managed components you can use:
- Queue: SQS / PubSub / Azure Queue
- Storage: S3 / GCS / Blob
- Compute: Lambda / Cloud Run / Fargate
- DB for state: DynamoDB / Firestore / Postgres
- Orchestration: Step Functions / Workflows / Temporal Cloud
Key engineering requirements at million-page scale
Must-haves
- URL deduplication
- Retry logic
- Per-domain rate limiting
- Backoff on 429/403/5xx
- Canonicalization of URLs
- Content fingerprinting
- Checkpointing
- Monitoring and cost controls
Nice-to-haves
- robots.txt awareness
- JS rendering only when needed
- language/content filters
- automatic sitemap discovery
- proxy rotation if allowed and necessary
What I’d recommend based on your goal
If you want the fastest path
Use Apify or Zyte.
If you want custom control but no servers
Use Cloud Run + a managed queue + object storage.
If you just need data at scale
Use Common Crawl or a commercial dataset.
Important note
Crawling millions of pages can easily run into:
- site terms of service violations
- copyright issues
- anti-bot protections
- privacy issues
So it’s best to crawl:
- your own sites
- sites that allow it
- public datasets
- APIs where available
If you want, I can give you:
- a cheap architecture on AWS/GCP,
- a tool-by-tool comparison, or
- a reference crawler design for million-page scale.