Prompt
How do I monitor website changes with an automated crawl?
Latest observation
To monitor website changes with an automated crawl, you typically set up a job that regularly visits pages, extracts the content you care about, and compares it to the previous version. Here’s a practical approach:
1) Define what you want to monitor
Decide whether you need to track:
- Full page HTML
- Visible text only
- Specific elements like prices, headlines, product availability, or metadata
- A list of pages or a site section
Monitoring a smaller, targeted part of the page is usually more reliable than diffing the entire HTML.
2) Crawl on a schedule
Run your crawler at a fixed interval:
- Every hour
- Daily
- Weekly
- On-demand for important pages
Use:
- Cron jobs
- A task scheduler
- A serverless scheduled function
- A workflow tool like Airflow, Prefect, or GitHub Actions
3) Fetch pages consistently
Use the same method every time:
- Static pages:
requests/curl - JavaScript-heavy pages: Playwright, Puppeteer, or Selenium
Be sure to:
- Respect
robots.txt - Rate-limit requests
- Set a clear user agent
- Avoid overloading the site
4) Extract stable content
Before comparing, clean the page data:
- Remove navigation, ads, timestamps, and dynamic widgets
- Normalize whitespace
- Strip tracking query parameters
- Convert relative links if needed
Example: compare just the product title, price, and stock status instead of the entire page source.
5) Store snapshots
Save each crawl result with:
- URL
- Timestamp
- Extracted content
- Content hash
- Optional HTML snapshot or screenshot
A hash is useful for fast change detection:
- If hash is unchanged, no difference
- If changed, run a diff
6) Detect changes
Common techniques:
- Hash comparison: simplest
- Text diff: shows what changed
- DOM diff: more structured, useful for HTML elements
- Visual diff: compares screenshots for design/layout changes
For example:
- Hash the cleaned text
- If it differs from last crawl, generate a diff and alert
7) Alert when changes occur
Send notifications via:
- Slack/Teams
- Webhooks
- SMS
- Dashboard
Include:
- URL
- What changed
- Old vs new value
- Time detected
8) Handle dynamic and protected pages
If the site:
- Loads content via JavaScript → use a headless browser
- Requires login → use authenticated sessions securely
- Has anti-bot defenses → reduce frequency, add retries, and comply with terms
9) Example simple workflow
- Crawl page every 6 hours
- Extract main content or selected fields
- Normalize and hash the result
- Compare with previous hash
- If different, store diff and notify
10) Minimal Python example
Here’s a simple text-based monitor:
import requests
from bs4 import BeautifulSoup
import hashlib
import json
from datetime import datetime
URL = "https://example.com"
STATE_FILE = "state.json"
def get_clean_text(url):
r = requests.get(url, timeout=20, headers={"User-Agent": "ChangeMonitor/1.0"})
r.raise_for_status()
soup = BeautifulSoup(r.text, "html.parser")
# Remove script/style
for tag in soup(["script", "style", "noscript"]):
tag.decompose()
text = " ".join(soup.get_text(separator=" ").split())
return text
def hash_text(text):
return hashlib.sha256(text.encode("utf-8")).hexdigest()
def load_state():
try:
with open(STATE_FILE, "r") as f:
return json.load(f)
except FileNotFoundError:
return {}
def save_state(state):
with open(STATE_FILE, "w") as f:
json.dump(state, f, indent=2)
text = get_clean_text(URL)
current_hash = hash_text(text)
state = load_state()
previous_hash = state.get(URL, {}).get("hash")
if previous_hash != current_hash:
print(f"Change detected on {URL} at {datetime.utcnow().isoformat()}Z")
print("Previous hash:", previous_hash)
print("Current hash:", current_hash)
state[URL] = {
"hash": current_hash,
"checked_at": datetime.utcnow().isoformat() + "Z"
}
save_state(state)
11) Better production setup
For a robust monitor, add:
- Retry logic
- Logging
- Error handling
- Retry backoff
- Per-page extraction rules
- Persistent database storage
- Notifications and audit history
If you want, I can also help you with:
- A Python/Playwright crawler
- A Docker + cron setup
- A database-backed website change monitor
- A tool comparison (Scrapy vs Playwright vs commercial tools)
Brands