Prompt

How do I build a website monitoring crawler?

Technology / Cloud Infrastructure1 observationLast seen Oct 4, 2026

Latest observation

Oct 4, 2026OpenAI APIWeb search: off

To build a website monitoring crawler, you usually want three parts:

  1. Crawler – visits pages and collects data
  2. Monitor – compares current vs previous snapshots
  3. Alerting/reporting – notifies you when something changes

Here’s a practical way to approach it.


1) Define what you want to monitor

Be specific about the changes you care about:

  • Page content changes: text, headings, product prices, news updates
  • HTML structure changes: DOM element added/removed
  • Visual changes: screenshots, layout shifts
  • Availability: page returns 200/404/500, slow response
  • Links/SEO: broken links, meta tags, titles

This matters because the crawler design differs for each.


2) Decide between static crawling and browser automation

Static HTTP crawler

Use this if pages are mostly server-rendered HTML.

  • Faster
  • Simpler
  • Lower resource usage

Tools:

  • Python: requests, httpx, BeautifulSoup, lxml
  • Node.js: axios, cheerio

Browser-based crawler

Use this if content is loaded by JavaScript.

  • Slower
  • More resource-heavy
  • Better for modern SPAs

Tools:

  • Playwright
  • Puppeteer
  • Selenium

A common approach is:

  • Try static fetch first
  • Fall back to a browser only when needed

3) Design the crawl flow

Typical loop:

  1. Start with a list of seed URLs
  2. Fetch page
  3. Parse and normalize content
  4. Extract relevant fields
  5. Save snapshot
  6. Compare with previous snapshot
  7. Trigger alerts if changed
  8. Revisit on a schedule

4) Respect crawl rules

Important for legal and operational safety:

  • Check robots.txt
  • Set a clear User-Agent
  • Rate limit requests
  • Avoid excessive concurrent requests
  • Don’t crawl pages you’re not allowed to access

Also watch for:

  • Login-protected content
  • CAPTCHAs
  • Terms of service restrictions

5) Store snapshots for comparison

You’ll want to keep:

  • URL
  • Fetch time
  • HTTP status
  • Raw HTML or extracted text
  • Hash of content
  • Screenshot path if using browser
  • Metadata like title, canonical URL, response headers

A simple storage option:

  • PostgreSQL or SQLite for metadata
  • S3/local disk for raw HTML/screenshots

6) Normalize content before diffing

Web pages often change in unimportant ways, like:

  • timestamps
  • CSRF tokens
  • dynamic ad content
  • random IDs

To reduce noise:

  • Strip scripts/styles
  • Remove known dynamic elements
  • Extract only relevant selectors
  • Normalize whitespace
  • Sort attributes if needed

For example, monitor:

  • product name
  • price
  • availability
  • article body instead of the entire page HTML.

7) Detect changes

Common techniques:

Hash comparison

  • Normalize page content
  • Compute hash
  • Compare with previous hash

Good for:

  • simple “changed / not changed” detection

Text diff

  • Compare extracted text lines
  • Useful for readable alerts

Structural diff

  • Compare DOM trees or selected elements
  • Good for layout/markup monitoring

Visual diff

  • Compare screenshots pixel-by-pixel or with perceptual diff

Good for:

  • design regressions
  • layout changes

8) Add scheduling and retries

Use a scheduler:

  • Cron
  • Celery beat
  • APScheduler
  • Airflow
  • a queue worker system

Also handle:

  • timeouts
  • retries with exponential backoff
  • temporary network failures
  • redirects
  • HTTP 429 rate limits

9) Build alerting

When a change is detected, notify via:

  • email
  • Slack
  • Discord
  • webhook
  • SMS

Alert contents should include:

  • URL
  • what changed
  • old vs new values
  • time of change
  • optional screenshot/diff

10) A simple Python example

Here’s a minimal static crawler/monitor:

import requests
from bs4 import BeautifulSoup
import hashlib
import sqlite3
import time

DB = "monitor.db"

def init_db():
    conn = sqlite3.connect(DB)
    conn.execute("""
        CREATE TABLE IF NOT EXISTS pages (
            url TEXT PRIMARY KEY,
            last_hash TEXT,
            last_content TEXT,
            last_checked INTEGER
        )
    """)
    conn.commit()
    conn.close()

def normalize_html(html):
    soup = BeautifulSoup(html, "html.parser")

    for tag in soup(["script", "style", "noscript"]):
        tag.decompose()

    text = soup.get_text(separator=" ", strip=True)
    return " ".join(text.split())

def get_hash(text):
    return hashlib.sha256(text.encode("utf-8")).hexdigest()

def monitor(url):
    headers = {"User-Agent": "MyMonitorBot/1.0"}
    response = requests.get(url, headers=headers, timeout=15)
    response.raise_for_status()

    content = normalize_html(response.text)
    content_hash = get_hash(content)

    conn = sqlite3.connect(DB)
    cur = conn.cursor()

    cur.execute("SELECT last_hash, last_content FROM pages WHERE url = ?", (url,))
    row = cur.fetchone()

    if row and row[0] != content_hash:
        old_content = row[1]
        print(f"CHANGE DETECTED for {url}")
        print("Old:", old_content[:300])
        print("New:", content[:300])

    cur.execute("""
        INSERT INTO pages (url, last_hash, last_content, last_checked)
        VALUES (?, ?, ?, ?)
        ON CONFLICT(url) DO UPDATE SET
            last_hash = excluded.last_hash,
            last_content = excluded.last_content,
            last_checked = excluded.last_checked
    """, (url, content_hash, content, int(time.time())))

    conn.commit()
    conn.close()

if __name__ == "__main__":
    init_db()
    monitor("https://example.com")

This:

  • fetches the page
  • removes scripts/styles
  • extracts visible text
  • hashes it
  • compares against the previous version
  • stores the latest snapshot

11) If the page uses JavaScript

Use Playwright. Example in Python:

from playwright.sync_api import sync_playwright
import hashlib

def get_page_text(url):
    with sync_playwright() as p:
        browser = p.chromium.launch(headless=True)
        page = browser.new_page()
        page.goto(url, wait_until="networkidle", timeout=30000)
        text = page.locator("body").inner_text()
        browser.close()
        return " ".join(text.split())

text = get_page_text("https://example.com")
print(hashlib.sha256(text.encode()).hexdigest())

12) Production considerations

If you want this to scale, add:

  • a queue for URLs
  • worker pool for fetching
  • persistent storage
  • per-domain rate limits
  • deduplication
  • metrics and logging
  • alert suppression for repeated changes
  • retries and dead-letter queues

If you need a larger system, a good stack is:

  • API/backend: FastAPI or Node.js
  • Queue: Redis + Celery/RQ/BullMQ
  • Storage: PostgreSQL
  • Browser rendering: Playwright workers
  • Scheduler: cron/Celery beat
  • Alerts: Slack/webhooks/email

13) Best practice: monitor selectors, not whole pages

Instead of monitoring all page text, target the exact elements:

  • CSS selectors like .price, #availability, article
  • XPath if needed
  • custom extraction logic

This reduces false positives.

Example:

price = soup.select_one(".price").get_text(strip=True)
availability = soup.select_one("#stock-status").get_text(strip=True)

14) Common pitfalls

  • Monitoring entire HTML and getting noisy diffs
  • Ignoring JS-rendered content
  • Not handling login/cookies
  • Crawling too fast and getting blocked
  • Forgetting to normalize timestamps/random tokens
  • No alert deduplication
  • Not storing historical versions

15) Recommended MVP

Start with this MVP:

  • Input: list of URLs
  • Fetch: requests or Playwright
  • Extract: target selectors or visible text
  • Save: SQLite/Postgres
  • Compare: hash against last snapshot
  • Alert: email or Slack
  • Schedule: cron every N minutes

If you want, I can also give you:

  1. a complete Python project structure,
  2. a Node.js version, or
  3. a Playwright-based crawler that monitors specific CSS selectors.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.