Prompt
How do I detect page updates across a big site list?
Latest observation
To detect page updates across a big list of sites, you generally want a change-detection pipeline rather than simple polling or full re-downloads.
Practical approaches
1) Use HTTP freshness signals first
For each URL, check:
ETagLast-ModifiedCache-Control/Expiresif relevant
Use conditional requests:
If-None-Match: <etag>If-Modified-Since: <date>
If the server returns 304 Not Modified, you can skip parsing.
This is the cheapest and most scalable method when supported.
2) Compare content hashes when freshness headers aren’t reliable
If the site doesn’t support ETag/Last-Modified well:
- Fetch the page
- Normalize the content a bit if needed:
- remove dynamic timestamps
- ignore ad/script blocks
- canonicalize whitespace
- Compute a hash:
- SHA-256 for exact changes
- SimHash / MinHash if you want near-duplicate detection
Store the previous hash and compare on each crawl.
3) Prefer sitemaps and feeds when available
Many sites expose:
sitemap.xml/ sitemap indexes- RSS/Atom feeds
These can reduce the number of pages you need to check and often include last update metadata.
4) Use diffing only after a change is detected
When a page changes:
- Store the previous HTML
- Generate a diff against the new version
- Extract structured change signals:
- title changes
- price changes
- text content changes
- schema/markup changes
This keeps the expensive work only for changed pages.
Scaling to a big site list
Recommended architecture
-
URL inventory
- Store all URLs, crawl frequency, priority, last seen status.
-
Scheduler
- Assign crawl intervals based on importance and change rate.
- Example: high-value pages hourly, stable pages daily/weekly.
-
Fetcher
- Send conditional GETs.
- Respect robots.txt and rate limits.
- Retry with backoff.
-
Change detector
- Check status code, headers, hash, or normalized diff.
-
Storage
- Save latest metadata and prior snapshots.
-
Alerting / downstream processing
- Trigger notifications or re-indexing when changes happen.
Efficiency tips
- Back off on unchanged pages: crawl less often if unchanged for long periods.
- Prioritize frequently changing URLs.
- Use concurrency carefully:
- async requests
- per-domain rate limiting
- Avoid full renders unless needed:
- use HTML fetch first
- headless browser only for JS-heavy pages
- Respect robots.txt and legal constraints.
If pages are JS-rendered
If the content changes after JavaScript runs:
- Use Playwright/Puppeteer only for pages that need it.
- Prefer extracting API endpoints if available.
- Compare rendered DOM or key data rather than raw HTML.
A simple rule of thumb
For a large list:
- First choice: conditional GET with
ETag/Last-Modified - Second choice: hash normalized content
- Third choice: diff only on change
- At scale: schedule adaptively and crawl by priority
If you want, I can also give you:
- a Python example for change detection,
- a database schema for tracking page versions, or
- a system design for crawling thousands/millions of URLs.
Brands