Prompt
I want to crawl a big list of sites automatically and keep revisiting them for updates, what should I use?
Latest observation
Oct 4, 2026
OpenAI APIWeb search: off
If you want to crawl lots of sites automatically and revisit them for changes, the best choice depends on whether you need simple page checks, full content extraction, or large-scale crawling.
Good options
1) Scrapy — best general-purpose crawler
Use this if you want to:
- crawl many sites
- follow links
- schedule repeat crawls
- extract structured data
- run at scale with queues/schedulers
Why it fits:
- built for crawling, not just single-page scraping
- efficient and fast
- easy to add rules for revisit frequency
- integrates well with databases, Redis, Celery, etc.
2) Playwright or Puppeteer — if sites are JavaScript-heavy
Use this if pages load content dynamically in the browser.
Why:
- handles modern JS-heavy sites
- can take screenshots, wait for selectors, interact with pages
Downside:
- slower and heavier than Scrapy
- more expensive at scale
A common setup is:
- Scrapy for most pages
- Playwright only for sites that need JavaScript
3) Apache Nutch / Heritrix — for very large-scale web crawling
Use these if you’re doing:
- huge-scale crawling
- web archiving
- broad crawling infrastructure
Why:
- designed for large crawls
- good for distributed setups
Downside:
- more operational complexity
- less convenient for custom app logic than Scrapy
4) Change detection / monitoring tools — if you only care about updates
If your goal is mostly “tell me when these pages change,” consider:
- a custom crawler that stores hashes/content snapshots
- specialized monitoring platforms
- RSS/Atom feeds where available
This is often better than fully crawling every page repeatedly.
What I’d recommend
For most use cases:
- Scrapy for crawling and revisiting
- PostgreSQL/MySQL to store pages, hashes, timestamps, and change history
- Redis + a scheduler to manage revisit intervals
- Playwright for JS-heavy sites only
Important things to plan for
- robots.txt compliance
- rate limiting
- retry/backoff
- deduplication
- change detection using hashes, normalized text, or DOM diffs
- per-site revisit policies based on how often content changes
Simple architecture
- Seed list of sites/URLs
- Crawl and fetch pages
- Extract canonical URL + content
- Normalize content
- Store hash and snapshot
- Revisit on schedule
- Compare hashes/content to detect updates
- Trigger alerts or downstream processing
My short answer
If you want one thing to start with: use Scrapy.
If many pages are JS-rendered: Scrapy + Playwright.
If you want, I can also suggest:
- a Python stack
- a Node.js stack
- or a production architecture for revisiting millions of URLs.