Prompt
How do I crawl many domains without maintaining servers?
Latest observation
If you want to crawl many domains without running your own servers, the best approach is usually to use a serverless / managed crawling setup or a distributed scraping API. The right choice depends on scale, frequency, and how “browser-heavy” the sites are.
Good options
1) Use a managed crawling platform
These services handle scheduling, proxies, retries, rendering, and storage for you.
Common patterns:
- Submit a list of URLs/domains
- They crawl pages, follow links, extract data
- You pull results via API or export
Examples of the type of tools to look for:
- Apify
- Zyte
- Bright Data Web Unlocker / Scraping Browser
- Diffbot
- Common Crawl if you only need existing web data
Best when:
- You need scale
- You don’t want ops
- You’re okay paying for convenience
2) Use serverless functions + queue/storage
If you want more control but still no servers:
- Use AWS Lambda, Cloud Run, Azure Functions, or Vercel/Cloudflare Workers for lightweight fetching
- Use a queue like SQS / PubSub / Redis Queue / Cloud Tasks
- Store results in S3 / GCS / Blob Storage / database
- Schedule with cron or managed schedulers
Typical architecture:
- Seed domain list in storage
- Queue crawl jobs
- Serverless worker fetches a page
- Extract links/content
- Enqueue discovered URLs
- Deduplicate and limit by domain rules
Best when:
- Pages are mostly static
- You want custom logic
- You’re comfortable building the pipeline
Caution:
- Pure serverless is awkward for long crawls, heavy crawling, or headless browser use
- Execution time limits can bite you
3) Use a scraping browser / browserless API for JS-heavy sites
If sites need JavaScript rendering, use a hosted browser service rather than running Puppeteer yourself.
Look for:
- Browserless
- Zyte Browser
- Playwright/Puppeteer hosted services
You send a URL, they return rendered HTML or screenshots.
Best when:
- The sites are dynamic
- You need login/session handling
- You don’t want to manage Chromium infrastructure
4) Use existing web datasets when possible
If you only need broad crawl data, consider:
- Common Crawl
- Search engine indexes / APIs
- Public data dumps
Best when:
- Freshness is less critical
- You need many domains but not real-time crawling
Practical architecture for many domains
A simple low-maintenance pattern:
- Input: list of root domains
- Frontier: managed queue
- Workers: serverless fetchers or managed crawl jobs
- Dedup: hash normalized URLs
- Politeness: per-domain rate limits and robots.txt awareness
- Storage: object storage + search index or database
- Monitoring: alerts on failures, crawl depth, latency
Key things to handle
To crawl responsibly and reliably:
- Respect robots.txt
- Add rate limits per domain
- Identify your crawler with a clear User-Agent
- Deduplicate URLs
- Normalize URLs to avoid loops
- Handle redirects, retries, and timeouts
- Separate static fetches from JS-rendered fetches
My recommendation
If you want the fastest path with minimal ops:
- Use a managed crawling service if budget allows
- Use serverless + queue if you want custom control and low maintenance
- Use a hosted browser API for JavaScript-heavy websites
If you tell me:
- how many domains/pages,
- whether pages are static or JS-heavy,
- how fresh the data needs to be,
- your cloud preference,
…I can suggest a concrete stack and architecture.