Prompt
How to monitor website changes daily with a crawler
Latest observation
Oct 4, 2026
OpenAI APIWeb search: off
You can monitor website changes daily with a crawler by following this setup:
1) Define what to watch
Decide which pages or sections matter:
- Product pages
- Pricing pages
- News/blog pages
- Legal/terms pages
- Specific HTML elements like titles, prices, dates, or text blocks
2) Build or use a crawler
A crawler should:
- Visit the target URL(s)
- Extract the content you care about
- Save a snapshot of the content
- Compare today’s snapshot to yesterday’s
You can do this with:
- Python:
requests,BeautifulSoup,scrapy,playwright - Node.js:
axios,cheerio,puppeteer - No-code tools: Visualping, Distill.io, ChangeTower
- SEO/site crawlers: Screaming Frog with scheduled runs
3) Normalize the page before comparing
Web pages often change in unimportant ways, so clean the content first:
- Remove scripts, ads, timestamps, random IDs, tracking params
- Extract only the relevant HTML section
- Normalize whitespace
- Sort attributes if needed
- Convert relative links to absolute if useful
4) Store snapshots
Save each daily crawl result in:
- A database
- JSON files
- Plain text files
- Cloud storage
Keep:
- URL
- Crawl time
- Extracted content
- Hash of the content
- Optional screenshot or HTML snapshot
5) Compare changes
Use one of these:
- Hash comparison: quick detection
- Diff comparison: shows exact text changes
- DOM comparison: more accurate for structured pages
- Screenshot comparison: useful for visual changes
6) Schedule daily runs
Use:
- Cron on Linux/macOS
- Task Scheduler on Windows
- GitHub Actions on a schedule
- Cloud functions / serverless cron
- Kubernetes CronJobs
Example cron:
0 6 * * * /usr/bin/python3 /path/to/crawler.py
7) Send alerts
When a change is detected, notify via:
- Slack
- Discord
- SMS
- Webhooks
Include:
- URL
- What changed
- Old vs new text
- Timestamp
8) Handle dynamic websites
If content loads via JavaScript, use:
- Playwright
- Puppeteer
- Selenium
These render the page before extraction.
9) Avoid common issues
- Respect
robots.txtand website terms - Rate-limit requests
- Use retries and timeouts
- Handle login sessions if needed
- Watch for false positives from ads, rotating banners, or timestamps
Simple workflow
- Crawl page daily
- Extract target text/section
- Generate hash or diff
- Compare with yesterday
- Alert on change
Minimal Python example
import requests
from bs4 import BeautifulSoup
import hashlib
import os
url = "https://example.com"
html = requests.get(url, timeout=10).text
soup = BeautifulSoup(html, "html.parser")
# Example: extract body text
text = soup.get_text(separator=" ", strip=True)
current_hash = hashlib.sha256(text.encode()).hexdigest()
prev_hash_file = "prev_hash.txt"
prev_hash = open(prev_hash_file).read().strip() if os.path.exists(prev_hash_file) else None
if prev_hash != current_hash:
print("Change detected!")
with open("current_snapshot.txt", "w") as f:
f.write(text)
with open(prev_hash_file, "w") as f:
f.write(current_hash)
Best practice
If you want reliable monitoring, combine:
- Targeted extraction
- Content normalization
- Scheduled crawling
- Diff-based alerts
If you want, I can also give you:
- a full Python crawler script,
- a Playwright version for JavaScript-heavy sites, or
- a production architecture for monitoring many pages daily.