Prompt
I'm building a website monitoring tool that needs scheduled crawls across many domains, what should I use?
Latest observation
Oct 4, 2026
OpenAI APIWeb search: off
For scheduled crawls across many domains, you generally want a crawler framework + job scheduler + distributed workers setup, rather than a single scraping library.
Good options
1. Scrapy + a scheduler
Best if you’re building in Python and want flexibility.
- Scrapy: strong crawling framework
- Cron / Celery / Airflow / APScheduler: for scheduling
- Redis / RabbitMQ + workers: for distributing jobs
- Good for:
- many domains
- recurring crawls
- custom extraction logic
- rate limiting and politeness controls
Common stack:
Scrapy + Celery + Redis + PostgreSQL
2. Apify
Good if you want a managed platform and don’t want to build infrastructure.
- Handles crawling, scheduling, queueing, retries
- Can run custom crawlers
- Good observability and scaling
- Useful if you want faster time to market
3. Playwright / Puppeteer + scheduler
Use this if sites are heavily JavaScript-rendered.
- Great for modern web apps
- More expensive than HTML crawling
- Pair with a scheduler and queue system
- Often used alongside a simpler HTTP crawler for most sites
4. Airflow / Temporal / Prefect
If scheduling and orchestration are the hardest part.
- Not crawlers themselves
- Best for managing lots of crawl workflows
- Useful when you have:
- retries
- dependencies
- per-domain schedules
- ETL pipelines after crawling
What I’d recommend
If you’re building this yourself:
- Crawler: Scrapy
- Scheduler/orchestration: Celery or Airflow
- Queue: Redis or RabbitMQ
- Storage: PostgreSQL + object storage for raw HTML
- JS-heavy fallback: Playwright
This is a solid architecture for:
- many domains
- recurring schedules
- retries/failures
- scaling horizontally
Important features to support
For website monitoring, make sure your tool includes:
- per-domain crawl frequency
- robots.txt compliance
- rate limiting / concurrency limits
- change detection
- retries and backoff
- deduplication
- alerting on detected changes
- proxy support if needed
- screenshot capture for visual monitoring
Simple rule of thumb
- Mostly static sites: Scrapy
- JS-heavy sites: Playwright
- Need managed service: Apify
- Need complex workflow scheduling: Airflow/Temporal
- Need easiest Python automation: Celery + Scrapy
If you want, I can suggest a specific architecture based on your scale, budget, and whether the sites are mostly static or JavaScript-heavy.