Prompt

How do I monitor website changes with an automated crawl?

Technology / Cloud Infrastructure1 observationLast seen Oct 4, 2026

Latest observation

Oct 4, 2026OpenAI APIWeb search: off

To monitor website changes with an automated crawl, you typically set up a job that regularly visits pages, extracts the content you care about, and compares it to the previous version. Here’s a practical approach:

1) Define what you want to monitor

Decide whether you need to track:

  • Full page HTML
  • Visible text only
  • Specific elements like prices, headlines, product availability, or metadata
  • A list of pages or a site section

Monitoring a smaller, targeted part of the page is usually more reliable than diffing the entire HTML.

2) Crawl on a schedule

Run your crawler at a fixed interval:

  • Every hour
  • Daily
  • Weekly
  • On-demand for important pages

Use:

  • Cron jobs
  • A task scheduler
  • A serverless scheduled function
  • A workflow tool like Airflow, Prefect, or GitHub Actions

3) Fetch pages consistently

Use the same method every time:

  • Static pages: requests/curl
  • JavaScript-heavy pages: Playwright, Puppeteer, or Selenium

Be sure to:

  • Respect robots.txt
  • Rate-limit requests
  • Set a clear user agent
  • Avoid overloading the site

4) Extract stable content

Before comparing, clean the page data:

  • Remove navigation, ads, timestamps, and dynamic widgets
  • Normalize whitespace
  • Strip tracking query parameters
  • Convert relative links if needed

Example: compare just the product title, price, and stock status instead of the entire page source.

5) Store snapshots

Save each crawl result with:

  • URL
  • Timestamp
  • Extracted content
  • Content hash
  • Optional HTML snapshot or screenshot

A hash is useful for fast change detection:

  • If hash is unchanged, no difference
  • If changed, run a diff

6) Detect changes

Common techniques:

  • Hash comparison: simplest
  • Text diff: shows what changed
  • DOM diff: more structured, useful for HTML elements
  • Visual diff: compares screenshots for design/layout changes

For example:

  • Hash the cleaned text
  • If it differs from last crawl, generate a diff and alert

7) Alert when changes occur

Send notifications via:

  • Email
  • Slack/Teams
  • Webhooks
  • SMS
  • Dashboard

Include:

  • URL
  • What changed
  • Old vs new value
  • Time detected

8) Handle dynamic and protected pages

If the site:

  • Loads content via JavaScript → use a headless browser
  • Requires login → use authenticated sessions securely
  • Has anti-bot defenses → reduce frequency, add retries, and comply with terms

9) Example simple workflow

  1. Crawl page every 6 hours
  2. Extract main content or selected fields
  3. Normalize and hash the result
  4. Compare with previous hash
  5. If different, store diff and notify

10) Minimal Python example

Here’s a simple text-based monitor:

import requests
from bs4 import BeautifulSoup
import hashlib
import json
from datetime import datetime

URL = "https://example.com"
STATE_FILE = "state.json"

def get_clean_text(url):
    r = requests.get(url, timeout=20, headers={"User-Agent": "ChangeMonitor/1.0"})
    r.raise_for_status()
    soup = BeautifulSoup(r.text, "html.parser")

    # Remove script/style
    for tag in soup(["script", "style", "noscript"]):
        tag.decompose()

    text = " ".join(soup.get_text(separator=" ").split())
    return text

def hash_text(text):
    return hashlib.sha256(text.encode("utf-8")).hexdigest()

def load_state():
    try:
        with open(STATE_FILE, "r") as f:
            return json.load(f)
    except FileNotFoundError:
        return {}

def save_state(state):
    with open(STATE_FILE, "w") as f:
        json.dump(state, f, indent=2)

text = get_clean_text(URL)
current_hash = hash_text(text)
state = load_state()

previous_hash = state.get(URL, {}).get("hash")
if previous_hash != current_hash:
    print(f"Change detected on {URL} at {datetime.utcnow().isoformat()}Z")
    print("Previous hash:", previous_hash)
    print("Current hash:", current_hash)

state[URL] = {
    "hash": current_hash,
    "checked_at": datetime.utcnow().isoformat() + "Z"
}
save_state(state)

11) Better production setup

For a robust monitor, add:

  • Retry logic
  • Logging
  • Error handling
  • Retry backoff
  • Per-page extraction rules
  • Persistent database storage
  • Notifications and audit history

If you want, I can also help you with:

  • A Python/Playwright crawler
  • A Docker + cron setup
  • A database-backed website change monitor
  • A tool comparison (Scrapy vs Playwright vs commercial tools)

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.