Prompt
How do I build a crawler that revisits pages for changes?
Latest observation
To build a crawler that revisits pages and detects changes, design it as a scheduled fetch + comparison + storage system.
Core idea
For each URL, you:
- Fetch the page
- Normalize the content
- Compare it to the previously stored version
- Store a new snapshot if it changed
- Schedule the next revisit based on your policy
1) Store metadata per URL
Keep a record like:
urllast_fetched_atlast_changed_atetag/last_modifiedheaders if availablecontent_hashstatus_codefetch_intervalpriorityfailure_count
This lets you avoid unnecessary full downloads and choose revisit frequency.
2) Use HTTP caching headers when possible
When revisiting, send conditional requests:
If-None-Match: <etag>If-Modified-Since: <last_modified>
If the server returns 304 Not Modified, you know the page hasn’t changed and you can skip further processing.
This is much cheaper than downloading and diffing the whole page.
3) Normalize content before hashing/comparing
Raw HTML changes often for irrelevant reasons, such as:
- timestamps
- ad markup
- rotating scripts
- tracking tokens
- whitespace or attribute order
To reduce false positives:
- strip scripts/styles if not needed
- extract the main content
- canonicalize whitespace
- remove known dynamic regions
- optionally render JS pages and compare the rendered DOM/text
Then compute a hash, such as:
- SHA-256 of normalized text
- SimHash / MinHash if you want near-duplicate detection
4) Decide what “change” means
You may care about different levels of change:
- Any HTML change
- Visible text change
- Structured field change (e.g., price, product availability)
- Semantic change (e.g., meaningfully different content)
If you’re tracking specific data, extract fields into structured records and compare those instead of the whole page.
Example:
- product page: compare
price,stock,title - news page: compare
headline,body text - job listing: compare
salary,location,description
5) Revisit scheduling strategy
Not all pages should be revisited equally often.
Common approaches:
- Fixed interval: revisit every N minutes/hours/days
- Adaptive interval:
- pages that change often → revisit more often
- pages that rarely change → revisit less often
- Priority-based:
- high-value pages first
- newly discovered pages more frequently at first
A simple adaptive rule:
- if changed recently, halve the interval
- if unchanged several times, gradually increase interval
6) Add a queue + worker architecture
A practical crawler usually uses:
- URL frontier / queue: URLs to revisit
- Scheduler: decides when a URL is due
- Workers: fetch pages concurrently
- Storage: snapshots and metadata
- Diff engine: detects changes
This helps you scale and retry failures cleanly.
7) Handle politeness and reliability
Respect sites and avoid being blocked:
- obey
robots.txtwhere appropriate - rate-limit requests per domain
- use retries with exponential backoff
- identify your crawler with a clear User-Agent
- don’t hammer pages that return errors
8) Example flow
For each URL:
- Check if it’s due for revisiting
- Fetch with conditional headers if available
- If
304, record “unchanged” - If
200, normalize content - Hash and compare to last snapshot
- If changed:
- save the new version
- emit a change event/notification
- Update scheduling metadata
9) Simple implementation sketch
Pseudo-code:
def revisit(url_record):
headers = {}
if url_record.etag:
headers["If-None-Match"] = url_record.etag
if url_record.last_modified:
headers["If-Modified-Since"] = url_record.last_modified
resp = fetch(url_record.url, headers=headers)
if resp.status_code == 304:
url_record.last_fetched_at = now()
return "unchanged"
content = normalize(resp.text)
new_hash = sha256(content)
if new_hash != url_record.content_hash:
save_snapshot(url_record.url, content, now())
url_record.content_hash = new_hash
url_record.last_changed_at = now()
changed = True
else:
changed = False
url_record.etag = resp.headers.get("ETag")
url_record.last_modified = resp.headers.get("Last-Modified")
url_record.last_fetched_at = now()
update_schedule(url_record, changed)
return "changed" if changed else "unchanged"
10) Tools you can use
Depending on your needs:
- Requests / httpx for HTTP fetching
- BeautifulSoup / lxml for HTML parsing
- Playwright / Selenium for JavaScript-heavy pages
- Redis / RabbitMQ / Celery / Kafka for queues
- PostgreSQL / SQLite / S3 for metadata and snapshots
11) Common pitfalls
- Comparing raw HTML and getting lots of false positives
- Ignoring JS-rendered content
- Re-crawling too aggressively
- Not handling redirects, errors, and timeouts
- Storing too many full snapshots without compression/deduplication
12) Best practice architecture
A good production setup is:
- Scheduler: decides next revisit time
- Fetcher: gets content with caching headers
- Normalizer: extracts stable content
- Comparator: hashes/diffs content
- Snapshot store: keeps versions
- Alerting/reporting: notifies when changes happen
If you want, I can also give you:
- a Python example using
requests+BeautifulSoup, - a distributed design for large-scale crawling, or
- a change-detection strategy for JavaScript-heavy sites.