Prompt

How do I monitor websites for changes and feed the results into analytics?

Technology / Data infrastructure1 observationLast seen Sep 24, 2026

Latest observation

Sep 24, 2026OpenAI APIWeb search: off

A good way to do this is to build a small pipeline with four parts:

  1. Check the website
  2. Detect what changed
  3. Store the result
  4. Send the result to analytics

Here’s a practical setup.


1) Decide what “changes” means

Different sites need different monitoring methods:

  • Full page HTML changes: any DOM/content change
  • Specific text changes: price, title, stock status, article body
  • Visual changes: layout/image differences
  • Structured data changes: JSON, API responses, tables
  • Event changes: availability, redirects, status codes, metadata

If you only care about a few fields, it’s better to extract those fields instead of comparing full page HTML.


2) Fetch the site on a schedule

Use a cron job, scheduler, or workflow runner to poll pages regularly.

Common options:

  • Cron + Python/Node script
  • Airflow / Dagster / Prefect for larger pipelines
  • Serverless scheduled jobs: AWS Lambda, GCP Cloud Scheduler + Cloud Run, Azure Functions
  • Scraping platforms: Apify, Zyte, Browse AI, etc.

Be polite:

  • Respect robots.txt where appropriate
  • Use reasonable intervals
  • Add retries, backoff, and timeouts
  • Identify your crawler if needed

3) Detect changes

Simple approach: hash the content

  • Fetch page
  • Normalize content
  • Hash it
  • Compare with last stored hash

Example normalization:

  • Remove whitespace
  • Remove script/style tags
  • Strip dynamic tokens like timestamps or nonce values

Better approach: compare extracted fields

Example:

  • Product price: $19.99
  • Availability: In stock
  • Headline: New policy announced

This reduces noise.

For visual changes

  • Take screenshots
  • Compare images using diff tools
  • Useful for UI/layout monitoring

For structured pages

  • Parse HTML with BeautifulSoup/Cheerio
  • Or call underlying APIs if available
  • Compare JSON values directly

4) Store change history

Save every check or only deltas, depending on volume.

Typical fields:

  • URL
  • Timestamp
  • Status code
  • Extracted values
  • Content hash
  • Change detected flag
  • Old value / new value
  • Error info

Good storage options:

  • PostgreSQL / MySQL for relational history
  • MongoDB for flexible documents
  • S3/GCS for raw snapshots and screenshots
  • Redis for short-term state

A common pattern is:

  • Store the latest fingerprint
  • Store change events
  • Store periodic raw snapshots for audit/debugging

5) Feed the results into analytics

Once you detect a change, send an event to your analytics system.

Event schema example

{
  "event_name": "website_change_detected",
  "url": "https://example.com/product/123",
  "timestamp": "2026-09-24T12:00:00Z",
  "change_type": "price_change",
  "field": "price",
  "old_value": "19.99",
  "new_value": "17.99",
  "source": "crawler_v1",
  "site_id": "example"
}

Where to send it

  • GA4 / Google Analytics
  • Mixpanel
  • Amplitude
  • Segment
  • Snowflake / BigQuery / Databricks
  • Kafka / Kinesis / Pub/Sub if you have a data platform
  • Webhook to your internal API

If you use an event pipeline, the monitor can publish to a queue, and analytics consumers can subscribe.


6) Recommended architecture

A common production design:

Scheduler → Fetcher → Parser/Extractor → Diff engine → Event store → Analytics sink

Optional:

  • Queue between fetcher and diff engine for scale
  • Dead-letter queue for failures
  • Dashboard for monitoring jobs

7) Minimal Python example

This example checks a page, hashes cleaned text, and emits a change event.

import requests
from bs4 import BeautifulSoup
import hashlib
import json
from datetime import datetime

def clean_text(html):
    soup = BeautifulSoup(html, "html.parser")
    for tag in soup(["script", "style"]):
        tag.decompose()
    text = " ".join(soup.get_text(" ").split())
    return text

def fingerprint(text):
    return hashlib.sha256(text.encode("utf-8")).hexdigest()

url = "https://example.com"
html = requests.get(url, timeout=20).text
text = clean_text(html)
current_hash = fingerprint(text)

# load previous hash from DB/file/cache
previous_hash = None

if previous_hash != current_hash:
    event = {
        "event_name": "website_change_detected",
        "url": url,
        "timestamp": datetime.utcnow().isoformat() + "Z",
        "change_type": "content_change",
        "old_hash": previous_hash,
        "new_hash": current_hash,
    }
    print(json.dumps(event))

In a real system, replace previous_hash = None with a lookup from your database.


8) If you want analytics in a BI warehouse

A very effective pattern is:

  • Emit events to Kafka/PubSub/queue
  • Write them into BigQuery/Snowflake/Postgres
  • Build dashboards in Looker/Tableau/Power BI/Superset

Example metrics:

  • Changes per site per day
  • Average time between changes
  • Failure rate of checks
  • Price change frequency
  • Number of pages with content updates
  • Top sites by volatility

9) Practical tips

  • Use ETags and Last-Modified headers when available
  • Handle JavaScript-heavy pages with Playwright/Puppeteer
  • Normalize away noise like:
    • timestamps
    • ads
    • rotating banners
    • session IDs
  • Rate-limit your crawler
  • Log:
    • status codes
    • parse failures
    • timeouts
    • captcha blocks
  • Version your extraction rules so analytics can track parser changes

10) Best stack choices by complexity

Small project

  • Python
  • Requests + BeautifulSoup
  • Cron
  • SQLite/Postgres
  • Webhook or CSV to analytics

Medium project

  • Python/Node
  • Playwright for dynamic pages
  • Postgres
  • Queue (SQS/RabbitMQ)
  • Segment/BigQuery

Large project

  • Scheduler + workers
  • Kafka/PubSub
  • Playwright + parsers
  • Data warehouse
  • BI dashboards
  • Alerting/observability

If you want, I can also give you:

  1. a ready-to-run Python scraper + change detector,
  2. a database schema for change history, or
  3. a pipeline design for GA4 / BigQuery / Snowflake.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as Circuit. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.