Prompt
What should I use for recurring website monitoring crawls?
Latest observation
Oct 4, 2026
OpenAI APIWeb search: off
For recurring website monitoring crawls, the best choice is usually one of these, depending on what you’re monitoring:
1. If you want a simple, reliable production crawler
Use Python + Scrapy
- Great for recurring crawls
- Mature ecosystem
- Handles retries, throttling, and scheduling well
- Easy to store diffs and detect changes
2. If you need browser-rendered pages
Use Playwright
- Best for JavaScript-heavy sites
- More reliable than Selenium for modern web apps
- Good for monitoring dynamic content changes
A common setup is:
- Scrapy for normal HTML pages
- Playwright for pages that need JS rendering
3. If you want managed infrastructure
Use a hosted scraping platform like:
- Apify
- Bright Data
- Zyte
These are useful if you want:
- proxy rotation
- anti-bot handling
- scheduling
- less ops work
4. If you’re just doing lightweight periodic checks
Use:
- cron + requests/BeautifulSoup
- or cron + Playwright
This works well for a small number of sites and simple monitoring.
My practical recommendation
- Simple static sites:
Scrapy - JS-heavy sites:
Playwright - Large-scale recurring monitoring:
Scrapy + Playwright + scheduler + storage/diffing - Want minimal maintenance: managed service
Good monitoring stack
A solid production stack is often:
- Crawler: Scrapy or Playwright
- Scheduler: cron, Airflow, Celery, or a cloud scheduler
- Storage: PostgreSQL / S3 / Elasticsearch
- Change detection: hash diffs, DOM diffs, or normalized text comparison
- Alerts: email, Slack, webhook
If you want, I can suggest the best stack for your exact case based on:
- number of sites
- frequency
- JS rendering needs
- budget
- whether you need alerts or change detection