Prompt
How do I monitor websites for changes and feed the results into analytics?
Latest observation
A good way to do this is to build a small pipeline with four parts:
- Check the website
- Detect what changed
- Store the result
- Send the result to analytics
Here’s a practical setup.
1) Decide what “changes” means
Different sites need different monitoring methods:
- Full page HTML changes: any DOM/content change
- Specific text changes: price, title, stock status, article body
- Visual changes: layout/image differences
- Structured data changes: JSON, API responses, tables
- Event changes: availability, redirects, status codes, metadata
If you only care about a few fields, it’s better to extract those fields instead of comparing full page HTML.
2) Fetch the site on a schedule
Use a cron job, scheduler, or workflow runner to poll pages regularly.
Common options:
- Cron + Python/Node script
- Airflow / Dagster / Prefect for larger pipelines
- Serverless scheduled jobs: AWS Lambda, GCP Cloud Scheduler + Cloud Run, Azure Functions
- Scraping platforms: Apify, Zyte, Browse AI, etc.
Be polite:
- Respect robots.txt where appropriate
- Use reasonable intervals
- Add retries, backoff, and timeouts
- Identify your crawler if needed
3) Detect changes
Simple approach: hash the content
- Fetch page
- Normalize content
- Hash it
- Compare with last stored hash
Example normalization:
- Remove whitespace
- Remove script/style tags
- Strip dynamic tokens like timestamps or nonce values
Better approach: compare extracted fields
Example:
- Product price:
$19.99 - Availability:
In stock - Headline:
New policy announced
This reduces noise.
For visual changes
- Take screenshots
- Compare images using diff tools
- Useful for UI/layout monitoring
For structured pages
- Parse HTML with BeautifulSoup/Cheerio
- Or call underlying APIs if available
- Compare JSON values directly
4) Store change history
Save every check or only deltas, depending on volume.
Typical fields:
- URL
- Timestamp
- Status code
- Extracted values
- Content hash
- Change detected flag
- Old value / new value
- Error info
Good storage options:
- PostgreSQL / MySQL for relational history
- MongoDB for flexible documents
- S3/GCS for raw snapshots and screenshots
- Redis for short-term state
A common pattern is:
- Store the latest fingerprint
- Store change events
- Store periodic raw snapshots for audit/debugging
5) Feed the results into analytics
Once you detect a change, send an event to your analytics system.
Event schema example
{
"event_name": "website_change_detected",
"url": "https://example.com/product/123",
"timestamp": "2026-09-24T12:00:00Z",
"change_type": "price_change",
"field": "price",
"old_value": "19.99",
"new_value": "17.99",
"source": "crawler_v1",
"site_id": "example"
}
Where to send it
- GA4 / Google Analytics
- Mixpanel
- Amplitude
- Segment
- Snowflake / BigQuery / Databricks
- Kafka / Kinesis / Pub/Sub if you have a data platform
- Webhook to your internal API
If you use an event pipeline, the monitor can publish to a queue, and analytics consumers can subscribe.
6) Recommended architecture
A common production design:
Scheduler → Fetcher → Parser/Extractor → Diff engine → Event store → Analytics sink
Optional:
- Queue between fetcher and diff engine for scale
- Dead-letter queue for failures
- Dashboard for monitoring jobs
7) Minimal Python example
This example checks a page, hashes cleaned text, and emits a change event.
import requests
from bs4 import BeautifulSoup
import hashlib
import json
from datetime import datetime
def clean_text(html):
soup = BeautifulSoup(html, "html.parser")
for tag in soup(["script", "style"]):
tag.decompose()
text = " ".join(soup.get_text(" ").split())
return text
def fingerprint(text):
return hashlib.sha256(text.encode("utf-8")).hexdigest()
url = "https://example.com"
html = requests.get(url, timeout=20).text
text = clean_text(html)
current_hash = fingerprint(text)
# load previous hash from DB/file/cache
previous_hash = None
if previous_hash != current_hash:
event = {
"event_name": "website_change_detected",
"url": url,
"timestamp": datetime.utcnow().isoformat() + "Z",
"change_type": "content_change",
"old_hash": previous_hash,
"new_hash": current_hash,
}
print(json.dumps(event))
In a real system, replace previous_hash = None with a lookup from your database.
8) If you want analytics in a BI warehouse
A very effective pattern is:
- Emit events to Kafka/PubSub/queue
- Write them into BigQuery/Snowflake/Postgres
- Build dashboards in Looker/Tableau/Power BI/Superset
Example metrics:
- Changes per site per day
- Average time between changes
- Failure rate of checks
- Price change frequency
- Number of pages with content updates
- Top sites by volatility
9) Practical tips
- Use ETags and Last-Modified headers when available
- Handle JavaScript-heavy pages with Playwright/Puppeteer
- Normalize away noise like:
- timestamps
- ads
- rotating banners
- session IDs
- Rate-limit your crawler
- Log:
- status codes
- parse failures
- timeouts
- captcha blocks
- Version your extraction rules so analytics can track parser changes
10) Best stack choices by complexity
Small project
- Python
- Requests + BeautifulSoup
- Cron
- SQLite/Postgres
- Webhook or CSV to analytics
Medium project
- Python/Node
- Playwright for dynamic pages
- Postgres
- Queue (SQS/RabbitMQ)
- Segment/BigQuery
Large project
- Scheduler + workers
- Kafka/PubSub
- Playwright + parsers
- Data warehouse
- BI dashboards
- Alerting/observability
If you want, I can also give you:
- a ready-to-run Python scraper + change detector,
- a database schema for change history, or
- a pipeline design for GA4 / BigQuery / Snowflake.
Brands