Prompt
How do I detect changes on websites with a crawler?
Latest observation
Oct 4, 2026
OpenAI APIWeb search: off
To detect changes on websites with a crawler, the usual approach is:
1) Crawl the page and store a baseline
Fetch the page content and save a representation of it, such as:
- raw HTML
- extracted text
- a normalized DOM
- selected fields from structured data
2) Revisit on a schedule
Run the crawler again periodically and compare the new version to the saved baseline.
3) Compare in a meaningful way
Different comparison methods detect different kinds of changes:
Simple methods
- Hash the full HTML
- Fast and easy
- Detects any change at all
- Very sensitive to irrelevant changes like timestamps, ads, session IDs
Better methods
- Normalize before comparing
- Remove scripts, styles, timestamps, tracking parameters, random IDs
- Canonicalize whitespace and attribute ordering
- Compare extracted text
- Useful for content changes
- Compare DOM structure
- Useful when page layout or markup changes matter
- Field-level diff
- Best if you know what data you care about, e.g. title, price, stock status
4) Decide what counts as a change
Not every difference should trigger an alert. Common strategies:
- Ignore known dynamic sections
- Use thresholds for small text diffs
- Alert only when important fields change
- Track multiple change types:
- content change
- structural change
- visual change
5) For JavaScript-heavy sites, render the page
If content is loaded dynamically:
- use a headless browser like Playwright, Puppeteer, or Selenium
- wait for the page to finish loading
- then capture the rendered HTML or DOM
6) Optional: use visual diffing
For pages where appearance matters:
- take screenshots
- compare pixels or use perceptual hash methods
- useful for UI changes, but more expensive
Typical workflow
- Crawl page
- Render if needed
- Clean/normalize content
- Save snapshot
- Re-crawl later
- Compare snapshots
- Trigger alert if meaningful differences are found
Example approach in practice
If you are monitoring product pages:
- extract product name, price, availability
- store those fields in a database
- compare them on each crawl
- send alert if price or stock status changes
That is usually much more reliable than comparing the whole HTML.
Common pitfalls
- pages contain dynamic ads or rotating content
- timestamps make every crawl look different
- anti-bot protections block the crawler
- JavaScript renders data after initial page load
- the site changes structure often, breaking selectors
If you want, I can also show:
- a Python example for detecting page changes,
- a Playwright-based solution for JavaScript sites, or
- a robust architecture for monitoring thousands of URLs.