Prompt

How do I crawl a list of websites and keep them updated?

Technology / Cloud Infrastructure1 observationLast seen Oct 4, 2026

Latest observation

Oct 4, 2026OpenAI APIWeb search: off

To crawl a list of websites and keep them updated, you want a scheduled, incremental crawler rather than a one-time scraper.

1) Define what “updated” means

Decide whether you need:

  • New pages only
  • Changed content on known pages
  • Deleted pages
  • Metadata changes like title, price, date, etc.

This determines how often you crawl and what you store.

2) Store a canonical URL list

Keep a database/table with at least:

  • url
  • site_id
  • last_crawled_at
  • last_modified_at from the server, if available
  • etag if available
  • content_hash
  • status_code
  • crawl_frequency
  • next_crawl_at

This lets you crawl only when needed.

3) Use conditional requests first

For each URL, send:

  • If-Modified-Since
  • If-None-Match

If the server returns 304 Not Modified, you can skip downloading/parsing the page.

This is the easiest way to keep things updated efficiently.

4) Re-crawl based on priority

Not all sites/pages need the same frequency. Use rules like:

  • News pages: every 15 minutes
  • Product pages: every few hours
  • Static pages: every few days
  • Low-change pages: weekly

Adjust based on observed change rates.

5) Detect changes with hashes

After downloading content:

  • normalize it a bit
  • compute a hash
  • compare with the previous hash

If the hash changed, store a new version or update the record.

6) Respect robots.txt and rate limits

To avoid blocking and stay compliant:

  • check robots.txt
  • set a clear user agent
  • add delays per domain
  • limit concurrent requests per site
  • back off on errors and 429s

7) Keep a queue + scheduler

A typical architecture:

  • Scheduler selects URLs due for crawling
  • Queue holds crawl jobs
  • Workers fetch and parse pages
  • Database stores results and crawl state

This is much more scalable than a simple loop.

8) Discover new URLs if needed

If you want to find new pages on the same sites:

  • follow internal links
  • parse XML sitemaps
  • watch RSS/Atom feeds
  • crawl breadth-first with depth limits

9) Handle failures gracefully

Track:

  • timeouts
  • DNS errors
  • HTTP 4xx/5xx
  • redirects
  • parsing errors

Use retry logic with exponential backoff.

10) Typical implementation stack

Common choices:

  • Python
    • requests or httpx
    • BeautifulSoup / lxml
    • Scrapy for full crawling
    • Celery or cron for scheduling
  • Databases
    • PostgreSQL, SQLite for small jobs
    • Redis for queues
  • For dynamic sites
    • Playwright or Selenium if JavaScript rendering is required

Simple workflow

  1. Seed your website list
  2. Crawl each URL
  3. Save content hash and metadata
  4. On the next run, use conditional GET
  5. If changed, store new version
  6. Reschedule based on site/page change frequency

Example strategy

For each URL:

  • fetch headers first or use conditional GET
  • if modified, download page
  • extract desired data
  • compute hash
  • update database
  • set next crawl time

Important caution

If the sites are not yours, make sure your crawling is allowed by:

  • terms of service
  • robots.txt
  • legal/privacy requirements

If you want, I can also give you:

  1. a Python example using Scrapy
  2. a simple cron + requests implementation
  3. a database schema for tracking crawl state

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.