Prompt

How do I build a crawler that revisits pages for changes?

Technology / Cloud Infrastructure1 observationLast seen Oct 4, 2026

Latest observation

Oct 4, 2026OpenAI APIWeb search: off

To build a crawler that revisits pages and detects changes, design it as a scheduled fetch + comparison + storage system.

Core idea

For each URL, you:

  1. Fetch the page
  2. Normalize the content
  3. Compare it to the previously stored version
  4. Store a new snapshot if it changed
  5. Schedule the next revisit based on your policy

1) Store metadata per URL

Keep a record like:

  • url
  • last_fetched_at
  • last_changed_at
  • etag / last_modified headers if available
  • content_hash
  • status_code
  • fetch_interval
  • priority
  • failure_count

This lets you avoid unnecessary full downloads and choose revisit frequency.


2) Use HTTP caching headers when possible

When revisiting, send conditional requests:

  • If-None-Match: <etag>
  • If-Modified-Since: <last_modified>

If the server returns 304 Not Modified, you know the page hasn’t changed and you can skip further processing.

This is much cheaper than downloading and diffing the whole page.


3) Normalize content before hashing/comparing

Raw HTML changes often for irrelevant reasons, such as:

  • timestamps
  • ad markup
  • rotating scripts
  • tracking tokens
  • whitespace or attribute order

To reduce false positives:

  • strip scripts/styles if not needed
  • extract the main content
  • canonicalize whitespace
  • remove known dynamic regions
  • optionally render JS pages and compare the rendered DOM/text

Then compute a hash, such as:

  • SHA-256 of normalized text
  • SimHash / MinHash if you want near-duplicate detection

4) Decide what “change” means

You may care about different levels of change:

  • Any HTML change
  • Visible text change
  • Structured field change (e.g., price, product availability)
  • Semantic change (e.g., meaningfully different content)

If you’re tracking specific data, extract fields into structured records and compare those instead of the whole page.

Example:

  • product page: compare price, stock, title
  • news page: compare headline, body text
  • job listing: compare salary, location, description

5) Revisit scheduling strategy

Not all pages should be revisited equally often.

Common approaches:

  • Fixed interval: revisit every N minutes/hours/days
  • Adaptive interval:
    • pages that change often → revisit more often
    • pages that rarely change → revisit less often
  • Priority-based:
    • high-value pages first
    • newly discovered pages more frequently at first

A simple adaptive rule:

  • if changed recently, halve the interval
  • if unchanged several times, gradually increase interval

6) Add a queue + worker architecture

A practical crawler usually uses:

  • URL frontier / queue: URLs to revisit
  • Scheduler: decides when a URL is due
  • Workers: fetch pages concurrently
  • Storage: snapshots and metadata
  • Diff engine: detects changes

This helps you scale and retry failures cleanly.


7) Handle politeness and reliability

Respect sites and avoid being blocked:

  • obey robots.txt where appropriate
  • rate-limit requests per domain
  • use retries with exponential backoff
  • identify your crawler with a clear User-Agent
  • don’t hammer pages that return errors

8) Example flow

For each URL:

  1. Check if it’s due for revisiting
  2. Fetch with conditional headers if available
  3. If 304, record “unchanged”
  4. If 200, normalize content
  5. Hash and compare to last snapshot
  6. If changed:
    • save the new version
    • emit a change event/notification
  7. Update scheduling metadata

9) Simple implementation sketch

Pseudo-code:

def revisit(url_record):
    headers = {}
    if url_record.etag:
        headers["If-None-Match"] = url_record.etag
    if url_record.last_modified:
        headers["If-Modified-Since"] = url_record.last_modified

    resp = fetch(url_record.url, headers=headers)

    if resp.status_code == 304:
        url_record.last_fetched_at = now()
        return "unchanged"

    content = normalize(resp.text)
    new_hash = sha256(content)

    if new_hash != url_record.content_hash:
        save_snapshot(url_record.url, content, now())
        url_record.content_hash = new_hash
        url_record.last_changed_at = now()
        changed = True
    else:
        changed = False

    url_record.etag = resp.headers.get("ETag")
    url_record.last_modified = resp.headers.get("Last-Modified")
    url_record.last_fetched_at = now()

    update_schedule(url_record, changed)
    return "changed" if changed else "unchanged"

10) Tools you can use

Depending on your needs:

  • Requests / httpx for HTTP fetching
  • BeautifulSoup / lxml for HTML parsing
  • Playwright / Selenium for JavaScript-heavy pages
  • Redis / RabbitMQ / Celery / Kafka for queues
  • PostgreSQL / SQLite / S3 for metadata and snapshots

11) Common pitfalls

  • Comparing raw HTML and getting lots of false positives
  • Ignoring JS-rendered content
  • Re-crawling too aggressively
  • Not handling redirects, errors, and timeouts
  • Storing too many full snapshots without compression/deduplication

12) Best practice architecture

A good production setup is:

  • Scheduler: decides next revisit time
  • Fetcher: gets content with caching headers
  • Normalizer: extracts stable content
  • Comparator: hashes/diffs content
  • Snapshot store: keeps versions
  • Alerting/reporting: notifies when changes happen

If you want, I can also give you:

  1. a Python example using requests + BeautifulSoup,
  2. a distributed design for large-scale crawling, or
  3. a change-detection strategy for JavaScript-heavy sites.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.