Prompt

Help me design a large-scale website crawler with scheduling, dedupe, and export to a warehouse

Technology / Cloud Infrastructure1 observationLast seen Oct 4, 2026

Latest observation

Oct 4, 2026OpenAI APIWeb search: off

Here’s a practical design for a large-scale website crawler that supports scheduling, deduplication, and export to a warehouse. I’ll frame it as a production architecture you could actually build and operate.


1) What the system needs to do

At a high level, the crawler should:

  1. Accept crawl jobs
    • Crawl a domain, a sitemap, a list of seed URLs, or a recurring schedule.
  2. Discover URLs
    • Follow links, sitemaps, feeds, and optionally APIs.
  3. Fetch pages at scale
    • Respect robots.txt, rate limits, and per-host politeness.
  4. Deduplicate
    • Avoid recrawling identical content unnecessarily.
    • Avoid repeated URL fetches caused by canonicalization or URL variations.
  5. Store raw and processed data
    • Save HTML, metadata, extracted text, links, screenshots, hashes, etc.
  6. Export to a warehouse
    • Load structured crawl results into BigQuery/Snowflake/Redshift/Databricks.
  7. Schedule recrawls
    • Periodically revisit URLs based on freshness, change frequency, or priority.
  8. Scale horizontally
    • Support millions or billions of URLs with distributed workers.

2) Core architecture

A good pattern is a pipeline with durable queues.

Main components

  • Job API / Scheduler
    • Creates crawl jobs and schedules recrawls.
  • URL Frontier
    • Central system managing what URLs need to be crawled next.
  • Fetch Workers
    • Pull URLs from the frontier, fetch content, and emit results.
  • Parser / Extractor
    • Extracts links, metadata, content, canonical URLs, hashes.
  • Dedup Service
    • Determines whether a URL/content is new or already seen.
  • Storage Layer
    • Raw blobs in object storage, metadata in a database.
  • Warehouse Exporter
    • Periodically or continuously loads structured data into the warehouse.
  • Observability
    • Metrics, logs, tracing, crawl health dashboards.

3) Recommended data flow

Crawl lifecycle

  1. User creates a crawl job:
    • “Crawl example.com every 7 days”
  2. Scheduler places seed URLs into the frontier.
  3. Frontier selects URLs by:
    • priority
    • host politeness
    • crawl recency
    • allowed scope
  4. Fetch worker downloads page.
  5. Parser extracts:
    • title, headings, text
    • links
    • canonical URL
    • metadata
    • content hash
  6. Dedup logic decides:
    • should we store this as a new version?
    • should we suppress this URL because it’s duplicate?
  7. Raw page and metadata get stored.
  8. New links are normalized and pushed back to frontier.
  9. Structured records are exported to warehouse.

4) Scheduling design

Scheduling is usually more than “run every X hours.” You want adaptive recrawl scheduling.

Scheduling inputs

  • Crawl frequency configured per domain/URL pattern
  • Content change history
  • Page importance / priority
  • Host budget / politeness constraints
  • Failure history and retry state

Suggested scheduling model

Use a persistent crawl schedule table with fields like:

  • job_id
  • url
  • next_crawl_at
  • last_crawled_at
  • crawl_interval_seconds
  • priority
  • status
  • retry_count
  • etag, last_modified for conditional fetches
  • content_hash
  • discovered_from

The scheduler periodically scans for next_crawl_at <= now() and enqueues tasks into the frontier.

Adaptive recrawl logic

Adjust the next crawl time based on change frequency:

  • Pages that change often: recrawl sooner
  • Stable pages: recrawl less often
  • Error pages: back off exponentially
  • High-value pages: prioritize

A simple heuristic:

  • If content changed recently, halve interval
  • If unchanged for N crawls, increase interval
  • Cap with min/max interval bounds

5) URL deduplication

You need dedupe at multiple layers.

A. URL normalization dedupe

Before enqueueing:

  • lowercase scheme/host
  • remove default ports
  • sort query params if safe
  • strip fragments
  • normalize trailing slashes
  • remove tracking params if policy allows
  • punycode normalization
  • canonicalize path encoding

Store:

  • raw_url
  • normalized_url
  • url_fingerprint

B. URL frontier dedupe

Maintain a visited/enqueued set:

  • Redis set, Bloom filter, RocksDB, or a distributed KV store
  • Prevent repeated scheduling of same normalized URL

For very large scale:

  • Bloom filter for fast approximate membership
  • Back it with a persistent store for exactness

C. Content dedupe

Different URLs may return identical pages.

Use:

  • content_hash = hash(normalized content)
  • or simhash/minhash for near-duplicate detection

Store:

  • exact hash for exact duplicates
  • similarity hash for near-duplicates

D. Canonical dedupe

If page declares:

  • <link rel="canonical" href="..."> Then map content to canonical URL when appropriate.

Important: don’t blindly collapse everything to canonical; preserve source URL and discovered URL for lineage.


6) Fetching at scale

Fetch worker responsibilities

  • Get URL task
  • Check robots.txt and policy
  • Respect per-host rate limits
  • Fetch page
  • Handle redirects
  • Support HTTP compression
  • Capture headers and response metadata
  • Return structured result

Per-host politeness

Use a host-level scheduler:

  • one token bucket per host/domain
  • configurable concurrency per host
  • delay between requests

This prevents hammering a site and helps avoid bans.

Retry policy

Retry only on transient errors:

  • 429
  • 5xx
  • timeouts
  • network failures

Do not retry on:

  • 404
  • 410
  • robots disallow
  • permanent 4xx

Use exponential backoff with jitter.

Conditional fetches

Use:

  • If-None-Match: <etag>
  • If-Modified-Since: <last_modified>

This reduces bandwidth and helps identify unchanged pages.


7) Parsing and extraction

After fetching, run parser tasks.

Extract

  • title
  • meta description
  • headings
  • visible text
  • outgoing links
  • canonical URL
  • structured data (JSON-LD, microdata)
  • language
  • content type
  • response headers
  • render info if using a browser crawler

Optional browser rendering

For JS-heavy sites:

  • headless Chrome/Playwright worker pool
  • use selectively based on heuristics
  • expensive, so reserve for pages that need rendering

Typical strategy:

  • default to HTTP fetch
  • fall back to browser render when:
    • page is thin/empty HTML
    • known JS app
    • important domain in whitelist

8) Storage architecture

Use a split storage model.

Raw content storage

Store HTML, screenshots, and response payloads in object storage:

  • S3 / GCS / Azure Blob

Path example:

  • s3://crawl-raw/{job_id}/{domain}/{yyyy}/{mm}/{dd}/{url_fingerprint}.html.gz

Metadata store

Use a database for structured crawl state:

  • Postgres for moderate scale
  • Cassandra/ScyllaDB/DynamoDB for larger scale frontier/state
  • Elasticsearch/OpenSearch only for search, not primary state

Metadata tables/entities:

  • jobs
  • url_state
  • fetch_attempts
  • page_versions
  • link_graph
  • content_dedup_index

Why split storage?

  • Blob storage is cheap for large raw payloads
  • DB is good for indexed metadata and scheduling
  • Warehouse is good for analytics

9) Warehouse export design

The warehouse should receive clean, append-friendly tables.

Export options

  1. Batch export
    • Periodically write Parquet/JSONL to object storage
    • Load via COPY/LOAD jobs
  2. Streaming export
    • Publish events to Kafka/PubSub
    • Sink to warehouse via connector
  3. Hybrid
    • Stream important metadata, batch raw crawl facts

Recommended warehouse tables

crawl_jobs

  • job_id
  • source
  • crawl_type
  • created_at
  • schedule_type
  • status

urls

  • url_id
  • raw_url
  • normalized_url
  • domain
  • first_seen_at
  • last_seen_at

fetches

  • fetch_id
  • url_id
  • job_id
  • fetched_at
  • http_status
  • content_type
  • bytes_downloaded
  • response_time_ms
  • etag
  • last_modified
  • final_url

page_versions

  • page_version_id
  • url_id
  • fetch_id
  • content_hash
  • simhash
  • title
  • text_extracted
  • canonical_url
  • is_duplicate
  • changed_since_last

links

  • source_url_id
  • target_url
  • target_url_normalized
  • anchor_text
  • rel
  • discovered_at

This structure lets analysts query crawl behavior, freshness, and content change over time.


10) Dedup strategy in practice

Use a layered approach.

Exact duplicate detection

  • Compute hash on normalized content
  • Example:
    • strip whitespace
    • remove boilerplate if desired
    • lowercase if safe
  • Use SHA-256 or BLAKE3

Near-duplicate detection

Use one of:

  • SimHash
  • MinHash + LSH
  • document embeddings for semantic dedupe, if needed

Suggested policy

  • If exact hash matches existing version for same URL: mark as unchanged
  • If exact hash matches different URL: mark as duplicate content
  • If near-duplicate threshold exceeded: mark as near-duplicate but keep sample if needed

Important: keep source URL lineage even when content is deduped.


11) Frontier design

The frontier is the heart of the crawler.

Requirements

  • persistent
  • distributed
  • ordered by priority
  • respects host constraints
  • deduplicated
  • resumable after failure

Frontier data model

You can think of it as:

  • a priority queue of crawl tasks
  • grouped by host/domain
  • backed by durable storage

Suggested implementation pattern

  • Task queue: Kafka/SQS/PubSub
  • Host scheduler: DB/KV store
  • Dedup index: Bloom filter + persistent state
  • Worker lease system: tasks are leased, not permanently removed until acknowledged

Scheduling logic

When worker requests a task:

  1. select next eligible host
  2. check host crawl delay and concurrency
  3. pop next URL for that host
  4. lease it to worker with timeout
  5. on ack, mark done
  6. on timeout, requeue

12) Scaling approach

Horizontally scalable services

  • many fetch workers
  • many parser workers
  • many export workers

Bottlenecks to design around

  • frontier contention
  • per-host rate limits
  • database write amplification
  • object storage upload throughput
  • warehouse load latency

Techniques

  • partition by domain hash
  • batch writes
  • async uploads
  • caching robots.txt
  • use worker autoscaling based on queue depth and fetch latency
  • apply backpressure when warehouse/export lags

13) Reliability and correctness

Exactly-once is hard

At crawl scale, aim for at-least-once processing with idempotent writes.

Use idempotency keys like:

  • fetch_id
  • url_id + fetched_at_bucket
  • content_hash

Failure handling

  • worker crash: lease expires and task is retried
  • partial upload: object storage upload is atomic from system perspective if using temp key + rename/copy
  • duplicate task execution: dedupe at write time

Checkpointing

Keep:

  • last processed offset in queues
  • frontier state
  • export watermark
  • schedule watermark

14) Security and compliance

Include:

  • robots.txt compliance
  • allowlist/denylist
  • crawl rate policies
  • PII handling policy
  • secrets management
  • user-agent identification
  • opt-out handling
  • legal review for sensitive domains

If exporting content, consider:

  • data retention limits
  • redaction for sensitive fields
  • access control in warehouse

15) Suggested technology stack

One possible stack:

Control plane

  • API service: Go / Python / Java
  • Scheduler: Airflow / Temporal / custom service

Queue/frontier

  • Kafka for task stream
  • Redis or RocksDB for dedupe/cache
  • Postgres/Cassandra/DynamoDB for state

Workers

  • Fetcher: Go, Rust, or Python with asyncio
  • Browser rendering: Playwright/Chromium pool
  • Parser: Python or JVM

Storage

  • Raw content: S3/GCS
  • Metadata: Postgres for small/medium, Cassandra/DynamoDB for large
  • Search/index: OpenSearch if needed

Warehouse export

  • Parquet files in object storage + Snowpipe/COPY/BigQuery load
  • Or Kafka Connect / streaming sink

Orchestration/monitoring

  • Kubernetes
  • Prometheus + Grafana
  • OpenTelemetry
  • ELK / Loki

16) Minimal viable architecture

If you want to start simpler and evolve:

Phase 1

  • Scheduler writes URLs to Kafka
  • Workers fetch and store raw HTML in S3
  • Postgres stores metadata
  • Batch job writes parquet to warehouse

Phase 2

  • Add URL normalization and exact dedupe
  • Add robots and rate limiting
  • Add conditional GET
  • Add adaptive recrawl scheduling

Phase 3

  • Add near-duplicate detection
  • Add browser rendering
  • Add host-aware frontier
  • Add stronger distributed state

17) Example end-to-end workflow

  1. User submits example.com with weekly recrawl.
  2. Scheduler seeds homepage and sitemap URLs.
  3. Frontier dedupes and prioritizes URLs.
  4. Worker fetches /products.
  5. Parser extracts 200 links.
  6. Normalization removes tracking parameters.
  7. 120 new URLs enter frontier.
  8. Page hash matches a previous version, so content is unchanged.
  9. Metadata records new fetch but no new page version.
  10. Nightly export writes crawl facts to Snowflake.

18) Practical schema sketch

url_state

  • url_id
  • normalized_url
  • domain
  • first_seen_at
  • last_seen_at
  • last_crawled_at
  • next_crawl_at
  • visit_count
  • fail_count
  • last_http_status
  • content_hash
  • simhash
  • canonical_url

crawl_task

  • task_id
  • url_id
  • job_id
  • priority
  • lease_until
  • status
  • attempt_count

fetch_result

  • fetch_id
  • task_id
  • final_url
  • status_code
  • headers_json
  • raw_blob_uri
  • downloaded_at
  • duration_ms

extracted_link

  • source_url_id
  • target_url
  • target_normalized
  • anchor_text
  • nofollow
  • discovered_at

19) Key design choices to make early

Decide these up front:

  1. Scope
    • whole web, a set of domains, or enterprise sites?
  2. Freshness
    • how often should pages be revisited?
  3. Rendering
    • do you need JS rendering?
  4. Dedup policy
    • exact only, or near-duplicate too?
  5. Warehouse requirements
    • batch or streaming?
  6. Compliance constraints
    • robots, legal, PII, opt-outs
  7. Scale target
    • 1M URLs/day vs 1B URLs/day changes everything

20) Recommended architecture summary

If I had to choose a solid default:

  • Scheduler: Temporal or custom service
  • Frontier: Kafka + durable KV store
  • Dedup: URL normalization + Bloom filter + exact hash
  • Fetch: Go workers with per-host token buckets
  • Parsing: separate parser workers
  • Storage: S3 for raw, Postgres/Cassandra for metadata
  • Warehouse export: batch Parquet loads
  • Recrawl: adaptive scheduling based on change history
  • Monitoring: Prometheus/Grafana + structured logs

If you want, I can also provide any of these next:

  1. a diagram of the architecture,
  2. a database schema in SQL,
  3. a Kafka-based frontier design,
  4. a detailed scheduling algorithm,
  5. or a reference implementation in Python/Go.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as Circuit. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.