Prompt
How do I build a website monitoring crawler?
Latest observation
To build a website monitoring crawler, you usually want three parts:
- Crawler – visits pages and collects data
- Monitor – compares current vs previous snapshots
- Alerting/reporting – notifies you when something changes
Here’s a practical way to approach it.
1) Define what you want to monitor
Be specific about the changes you care about:
- Page content changes: text, headings, product prices, news updates
- HTML structure changes: DOM element added/removed
- Visual changes: screenshots, layout shifts
- Availability: page returns 200/404/500, slow response
- Links/SEO: broken links, meta tags, titles
This matters because the crawler design differs for each.
2) Decide between static crawling and browser automation
Static HTTP crawler
Use this if pages are mostly server-rendered HTML.
- Faster
- Simpler
- Lower resource usage
Tools:
- Python:
requests,httpx,BeautifulSoup,lxml - Node.js:
axios,cheerio
Browser-based crawler
Use this if content is loaded by JavaScript.
- Slower
- More resource-heavy
- Better for modern SPAs
Tools:
- Playwright
- Puppeteer
- Selenium
A common approach is:
- Try static fetch first
- Fall back to a browser only when needed
3) Design the crawl flow
Typical loop:
- Start with a list of seed URLs
- Fetch page
- Parse and normalize content
- Extract relevant fields
- Save snapshot
- Compare with previous snapshot
- Trigger alerts if changed
- Revisit on a schedule
4) Respect crawl rules
Important for legal and operational safety:
- Check
robots.txt - Set a clear
User-Agent - Rate limit requests
- Avoid excessive concurrent requests
- Don’t crawl pages you’re not allowed to access
Also watch for:
- Login-protected content
- CAPTCHAs
- Terms of service restrictions
5) Store snapshots for comparison
You’ll want to keep:
- URL
- Fetch time
- HTTP status
- Raw HTML or extracted text
- Hash of content
- Screenshot path if using browser
- Metadata like title, canonical URL, response headers
A simple storage option:
- PostgreSQL or SQLite for metadata
- S3/local disk for raw HTML/screenshots
6) Normalize content before diffing
Web pages often change in unimportant ways, like:
- timestamps
- CSRF tokens
- dynamic ad content
- random IDs
To reduce noise:
- Strip scripts/styles
- Remove known dynamic elements
- Extract only relevant selectors
- Normalize whitespace
- Sort attributes if needed
For example, monitor:
- product name
- price
- availability
- article body instead of the entire page HTML.
7) Detect changes
Common techniques:
Hash comparison
- Normalize page content
- Compute hash
- Compare with previous hash
Good for:
- simple “changed / not changed” detection
Text diff
- Compare extracted text lines
- Useful for readable alerts
Structural diff
- Compare DOM trees or selected elements
- Good for layout/markup monitoring
Visual diff
- Compare screenshots pixel-by-pixel or with perceptual diff
Good for:
- design regressions
- layout changes
8) Add scheduling and retries
Use a scheduler:
- Cron
- Celery beat
- APScheduler
- Airflow
- a queue worker system
Also handle:
- timeouts
- retries with exponential backoff
- temporary network failures
- redirects
- HTTP 429 rate limits
9) Build alerting
When a change is detected, notify via:
- Slack
- Discord
- webhook
- SMS
Alert contents should include:
- URL
- what changed
- old vs new values
- time of change
- optional screenshot/diff
10) A simple Python example
Here’s a minimal static crawler/monitor:
import requests
from bs4 import BeautifulSoup
import hashlib
import sqlite3
import time
DB = "monitor.db"
def init_db():
conn = sqlite3.connect(DB)
conn.execute("""
CREATE TABLE IF NOT EXISTS pages (
url TEXT PRIMARY KEY,
last_hash TEXT,
last_content TEXT,
last_checked INTEGER
)
""")
conn.commit()
conn.close()
def normalize_html(html):
soup = BeautifulSoup(html, "html.parser")
for tag in soup(["script", "style", "noscript"]):
tag.decompose()
text = soup.get_text(separator=" ", strip=True)
return " ".join(text.split())
def get_hash(text):
return hashlib.sha256(text.encode("utf-8")).hexdigest()
def monitor(url):
headers = {"User-Agent": "MyMonitorBot/1.0"}
response = requests.get(url, headers=headers, timeout=15)
response.raise_for_status()
content = normalize_html(response.text)
content_hash = get_hash(content)
conn = sqlite3.connect(DB)
cur = conn.cursor()
cur.execute("SELECT last_hash, last_content FROM pages WHERE url = ?", (url,))
row = cur.fetchone()
if row and row[0] != content_hash:
old_content = row[1]
print(f"CHANGE DETECTED for {url}")
print("Old:", old_content[:300])
print("New:", content[:300])
cur.execute("""
INSERT INTO pages (url, last_hash, last_content, last_checked)
VALUES (?, ?, ?, ?)
ON CONFLICT(url) DO UPDATE SET
last_hash = excluded.last_hash,
last_content = excluded.last_content,
last_checked = excluded.last_checked
""", (url, content_hash, content, int(time.time())))
conn.commit()
conn.close()
if __name__ == "__main__":
init_db()
monitor("https://example.com")
This:
- fetches the page
- removes scripts/styles
- extracts visible text
- hashes it
- compares against the previous version
- stores the latest snapshot
11) If the page uses JavaScript
Use Playwright. Example in Python:
from playwright.sync_api import sync_playwright
import hashlib
def get_page_text(url):
with sync_playwright() as p:
browser = p.chromium.launch(headless=True)
page = browser.new_page()
page.goto(url, wait_until="networkidle", timeout=30000)
text = page.locator("body").inner_text()
browser.close()
return " ".join(text.split())
text = get_page_text("https://example.com")
print(hashlib.sha256(text.encode()).hexdigest())
12) Production considerations
If you want this to scale, add:
- a queue for URLs
- worker pool for fetching
- persistent storage
- per-domain rate limits
- deduplication
- metrics and logging
- alert suppression for repeated changes
- retries and dead-letter queues
If you need a larger system, a good stack is:
- API/backend: FastAPI or Node.js
- Queue: Redis + Celery/RQ/BullMQ
- Storage: PostgreSQL
- Browser rendering: Playwright workers
- Scheduler: cron/Celery beat
- Alerts: Slack/webhooks/email
13) Best practice: monitor selectors, not whole pages
Instead of monitoring all page text, target the exact elements:
- CSS selectors like
.price,#availability,article - XPath if needed
- custom extraction logic
This reduces false positives.
Example:
price = soup.select_one(".price").get_text(strip=True)
availability = soup.select_one("#stock-status").get_text(strip=True)
14) Common pitfalls
- Monitoring entire HTML and getting noisy diffs
- Ignoring JS-rendered content
- Not handling login/cookies
- Crawling too fast and getting blocked
- Forgetting to normalize timestamps/random tokens
- No alert deduplication
- Not storing historical versions
15) Recommended MVP
Start with this MVP:
- Input: list of URLs
- Fetch: requests or Playwright
- Extract: target selectors or visible text
- Save: SQLite/Postgres
- Compare: hash against last snapshot
- Alert: email or Slack
- Schedule: cron every N minutes
If you want, I can also give you:
- a complete Python project structure,
- a Node.js version, or
- a Playwright-based crawler that monitors specific CSS selectors.