Prompt
Website change detection crawler
Latest observation
Oct 4, 2026
OpenAI APIWeb search: off
If you want a website change detection crawler, there are a few common approaches depending on what you need:
1. Simple page monitor
Checks a URL periodically and compares content to the previous version.
Best for: single pages, prices, announcements, docs.
How it works:
- Fetch HTML
- Normalize content
- Hash it
- Compare with last hash
- Alert if different
2. Headless browser crawler
Uses Playwright/Selenium to render JavaScript-heavy pages before comparing.
Best for: modern sites that load content dynamically.
3. DOM-aware diff crawler
Instead of comparing raw HTML, compare extracted text or specific selectors.
Best for: reducing false positives from ads, timestamps, or layout changes.
Basic Python example
Here’s a minimal crawler that detects changes to a page:
import requests
import hashlib
import os
URL = "https://example.com"
STATE_FILE = "page_hash.txt"
def fetch_page(url):
r = requests.get(url, timeout=20)
r.raise_for_status()
return r.text
def hash_content(content):
return hashlib.sha256(content.encode("utf-8")).hexdigest()
def load_previous_hash():
if os.path.exists(STATE_FILE):
with open(STATE_FILE, "r") as f:
return f.read().strip()
return None
def save_hash(h):
with open(STATE_FILE, "w") as f:
f.write(h)
def main():
content = fetch_page(URL)
current_hash = hash_content(content)
previous_hash = load_previous_hash()
if previous_hash and previous_hash != current_hash:
print("Change detected!")
elif not previous_hash:
print("First run, storing baseline.")
else:
print("No change.")
save_hash(current_hash)
if __name__ == "__main__":
main()
Better production features
- Respect
robots.txt - Rate limiting / backoff
- Store snapshots in a DB or S3
- Ignore dynamic noise with CSS selector extraction
- Send alerts via email/Slack/webhook
- Schedule with cron, Celery, Airflow, or a serverless job
If you want a scalable architecture
Typical components:
- Scheduler: triggers crawls
- Fetcher: HTTP or headless browser
- Normalizer: cleans content
- Comparator: detects changes
- Storage: hashes, snapshots, history
- Notifier: alerts on changes
If you want, I can generate:
- a full Python crawler project,
- a Playwright-based version, or
- a Dockerized service with alerts.