Prompt
How do I crawl a list of websites and keep them updated?
Latest observation
To crawl a list of websites and keep them updated, you want a scheduled, incremental crawler rather than a one-time scraper.
1) Define what “updated” means
Decide whether you need:
- New pages only
- Changed content on known pages
- Deleted pages
- Metadata changes like title, price, date, etc.
This determines how often you crawl and what you store.
2) Store a canonical URL list
Keep a database/table with at least:
urlsite_idlast_crawled_atlast_modified_atfrom the server, if availableetagif availablecontent_hashstatus_codecrawl_frequencynext_crawl_at
This lets you crawl only when needed.
3) Use conditional requests first
For each URL, send:
If-Modified-SinceIf-None-Match
If the server returns 304 Not Modified, you can skip downloading/parsing the page.
This is the easiest way to keep things updated efficiently.
4) Re-crawl based on priority
Not all sites/pages need the same frequency. Use rules like:
- News pages: every 15 minutes
- Product pages: every few hours
- Static pages: every few days
- Low-change pages: weekly
Adjust based on observed change rates.
5) Detect changes with hashes
After downloading content:
- normalize it a bit
- compute a hash
- compare with the previous hash
If the hash changed, store a new version or update the record.
6) Respect robots.txt and rate limits
To avoid blocking and stay compliant:
- check
robots.txt - set a clear user agent
- add delays per domain
- limit concurrent requests per site
- back off on errors and 429s
7) Keep a queue + scheduler
A typical architecture:
- Scheduler selects URLs due for crawling
- Queue holds crawl jobs
- Workers fetch and parse pages
- Database stores results and crawl state
This is much more scalable than a simple loop.
8) Discover new URLs if needed
If you want to find new pages on the same sites:
- follow internal links
- parse XML sitemaps
- watch RSS/Atom feeds
- crawl breadth-first with depth limits
9) Handle failures gracefully
Track:
- timeouts
- DNS errors
- HTTP 4xx/5xx
- redirects
- parsing errors
Use retry logic with exponential backoff.
10) Typical implementation stack
Common choices:
- Python
requestsorhttpxBeautifulSoup/lxmlScrapyfor full crawlingCeleryor cron for scheduling
- Databases
- PostgreSQL, SQLite for small jobs
- Redis for queues
- For dynamic sites
- Playwright or Selenium if JavaScript rendering is required
Simple workflow
- Seed your website list
- Crawl each URL
- Save content hash and metadata
- On the next run, use conditional GET
- If changed, store new version
- Reschedule based on site/page change frequency
Example strategy
For each URL:
- fetch headers first or use conditional GET
- if modified, download page
- extract desired data
- compute hash
- update database
- set next crawl time
Important caution
If the sites are not yours, make sure your crawling is allowed by:
- terms of service
- robots.txt
- legal/privacy requirements
If you want, I can also give you:
- a Python example using Scrapy
- a simple cron + requests implementation
- a database schema for tracking crawl state