Prompt

How do I detect page updates across a big site list?

Technology / Cloud Infrastructure1 observationLast seen Oct 4, 2026

Latest observation

Oct 4, 2026OpenAI APIWeb search: off

To detect page updates across a big list of sites, you generally want a change-detection pipeline rather than simple polling or full re-downloads.

Practical approaches

1) Use HTTP freshness signals first

For each URL, check:

  • ETag
  • Last-Modified
  • Cache-Control / Expires if relevant

Use conditional requests:

  • If-None-Match: <etag>
  • If-Modified-Since: <date>

If the server returns 304 Not Modified, you can skip parsing.
This is the cheapest and most scalable method when supported.


2) Compare content hashes when freshness headers aren’t reliable

If the site doesn’t support ETag/Last-Modified well:

  • Fetch the page
  • Normalize the content a bit if needed:
    • remove dynamic timestamps
    • ignore ad/script blocks
    • canonicalize whitespace
  • Compute a hash:
    • SHA-256 for exact changes
    • SimHash / MinHash if you want near-duplicate detection

Store the previous hash and compare on each crawl.


3) Prefer sitemaps and feeds when available

Many sites expose:

  • sitemap.xml / sitemap indexes
  • RSS/Atom feeds

These can reduce the number of pages you need to check and often include last update metadata.


4) Use diffing only after a change is detected

When a page changes:

  • Store the previous HTML
  • Generate a diff against the new version
  • Extract structured change signals:
    • title changes
    • price changes
    • text content changes
    • schema/markup changes

This keeps the expensive work only for changed pages.


Scaling to a big site list

Recommended architecture

  1. URL inventory

    • Store all URLs, crawl frequency, priority, last seen status.
  2. Scheduler

    • Assign crawl intervals based on importance and change rate.
    • Example: high-value pages hourly, stable pages daily/weekly.
  3. Fetcher

    • Send conditional GETs.
    • Respect robots.txt and rate limits.
    • Retry with backoff.
  4. Change detector

    • Check status code, headers, hash, or normalized diff.
  5. Storage

    • Save latest metadata and prior snapshots.
  6. Alerting / downstream processing

    • Trigger notifications or re-indexing when changes happen.

Efficiency tips

  • Back off on unchanged pages: crawl less often if unchanged for long periods.
  • Prioritize frequently changing URLs.
  • Use concurrency carefully:
    • async requests
    • per-domain rate limiting
  • Avoid full renders unless needed:
    • use HTML fetch first
    • headless browser only for JS-heavy pages
  • Respect robots.txt and legal constraints.

If pages are JS-rendered

If the content changes after JavaScript runs:

  • Use Playwright/Puppeteer only for pages that need it.
  • Prefer extracting API endpoints if available.
  • Compare rendered DOM or key data rather than raw HTML.

A simple rule of thumb

For a large list:

  • First choice: conditional GET with ETag / Last-Modified
  • Second choice: hash normalized content
  • Third choice: diff only on change
  • At scale: schedule adaptively and crawl by priority

If you want, I can also give you:

  1. a Python example for change detection,
  2. a database schema for tracking page versions, or
  3. a system design for crawling thousands/millions of URLs.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.