Prompt
Help me design a large-scale website crawler with scheduling, dedupe, and export to a warehouse
Latest observation
Here’s a practical design for a large-scale website crawler that supports scheduling, deduplication, and export to a warehouse. I’ll frame it as a production architecture you could actually build and operate.
1) What the system needs to do
At a high level, the crawler should:
- Accept crawl jobs
- Crawl a domain, a sitemap, a list of seed URLs, or a recurring schedule.
- Discover URLs
- Follow links, sitemaps, feeds, and optionally APIs.
- Fetch pages at scale
- Respect robots.txt, rate limits, and per-host politeness.
- Deduplicate
- Avoid recrawling identical content unnecessarily.
- Avoid repeated URL fetches caused by canonicalization or URL variations.
- Store raw and processed data
- Save HTML, metadata, extracted text, links, screenshots, hashes, etc.
- Export to a warehouse
- Load structured crawl results into BigQuery/Snowflake/Redshift/Databricks.
- Schedule recrawls
- Periodically revisit URLs based on freshness, change frequency, or priority.
- Scale horizontally
- Support millions or billions of URLs with distributed workers.
2) Core architecture
A good pattern is a pipeline with durable queues.
Main components
- Job API / Scheduler
- Creates crawl jobs and schedules recrawls.
- URL Frontier
- Central system managing what URLs need to be crawled next.
- Fetch Workers
- Pull URLs from the frontier, fetch content, and emit results.
- Parser / Extractor
- Extracts links, metadata, content, canonical URLs, hashes.
- Dedup Service
- Determines whether a URL/content is new or already seen.
- Storage Layer
- Raw blobs in object storage, metadata in a database.
- Warehouse Exporter
- Periodically or continuously loads structured data into the warehouse.
- Observability
- Metrics, logs, tracing, crawl health dashboards.
3) Recommended data flow
Crawl lifecycle
- User creates a crawl job:
- “Crawl example.com every 7 days”
- Scheduler places seed URLs into the frontier.
- Frontier selects URLs by:
- priority
- host politeness
- crawl recency
- allowed scope
- Fetch worker downloads page.
- Parser extracts:
- title, headings, text
- links
- canonical URL
- metadata
- content hash
- Dedup logic decides:
- should we store this as a new version?
- should we suppress this URL because it’s duplicate?
- Raw page and metadata get stored.
- New links are normalized and pushed back to frontier.
- Structured records are exported to warehouse.
4) Scheduling design
Scheduling is usually more than “run every X hours.” You want adaptive recrawl scheduling.
Scheduling inputs
- Crawl frequency configured per domain/URL pattern
- Content change history
- Page importance / priority
- Host budget / politeness constraints
- Failure history and retry state
Suggested scheduling model
Use a persistent crawl schedule table with fields like:
job_idurlnext_crawl_atlast_crawled_atcrawl_interval_secondsprioritystatusretry_countetag,last_modifiedfor conditional fetchescontent_hashdiscovered_from
The scheduler periodically scans for next_crawl_at <= now() and enqueues tasks into the frontier.
Adaptive recrawl logic
Adjust the next crawl time based on change frequency:
- Pages that change often: recrawl sooner
- Stable pages: recrawl less often
- Error pages: back off exponentially
- High-value pages: prioritize
A simple heuristic:
- If content changed recently, halve interval
- If unchanged for N crawls, increase interval
- Cap with min/max interval bounds
5) URL deduplication
You need dedupe at multiple layers.
A. URL normalization dedupe
Before enqueueing:
- lowercase scheme/host
- remove default ports
- sort query params if safe
- strip fragments
- normalize trailing slashes
- remove tracking params if policy allows
- punycode normalization
- canonicalize path encoding
Store:
raw_urlnormalized_urlurl_fingerprint
B. URL frontier dedupe
Maintain a visited/enqueued set:
- Redis set, Bloom filter, RocksDB, or a distributed KV store
- Prevent repeated scheduling of same normalized URL
For very large scale:
- Bloom filter for fast approximate membership
- Back it with a persistent store for exactness
C. Content dedupe
Different URLs may return identical pages.
Use:
content_hash = hash(normalized content)- or
simhash/minhashfor near-duplicate detection
Store:
- exact hash for exact duplicates
- similarity hash for near-duplicates
D. Canonical dedupe
If page declares:
<link rel="canonical" href="...">Then map content to canonical URL when appropriate.
Important: don’t blindly collapse everything to canonical; preserve source URL and discovered URL for lineage.
6) Fetching at scale
Fetch worker responsibilities
- Get URL task
- Check robots.txt and policy
- Respect per-host rate limits
- Fetch page
- Handle redirects
- Support HTTP compression
- Capture headers and response metadata
- Return structured result
Per-host politeness
Use a host-level scheduler:
- one token bucket per host/domain
- configurable concurrency per host
- delay between requests
This prevents hammering a site and helps avoid bans.
Retry policy
Retry only on transient errors:
- 429
- 5xx
- timeouts
- network failures
Do not retry on:
- 404
- 410
- robots disallow
- permanent 4xx
Use exponential backoff with jitter.
Conditional fetches
Use:
If-None-Match: <etag>If-Modified-Since: <last_modified>
This reduces bandwidth and helps identify unchanged pages.
7) Parsing and extraction
After fetching, run parser tasks.
Extract
- title
- meta description
- headings
- visible text
- outgoing links
- canonical URL
- structured data (JSON-LD, microdata)
- language
- content type
- response headers
- render info if using a browser crawler
Optional browser rendering
For JS-heavy sites:
- headless Chrome/Playwright worker pool
- use selectively based on heuristics
- expensive, so reserve for pages that need rendering
Typical strategy:
- default to HTTP fetch
- fall back to browser render when:
- page is thin/empty HTML
- known JS app
- important domain in whitelist
8) Storage architecture
Use a split storage model.
Raw content storage
Store HTML, screenshots, and response payloads in object storage:
- S3 / GCS / Azure Blob
Path example:
s3://crawl-raw/{job_id}/{domain}/{yyyy}/{mm}/{dd}/{url_fingerprint}.html.gz
Metadata store
Use a database for structured crawl state:
- Postgres for moderate scale
- Cassandra/ScyllaDB/DynamoDB for larger scale frontier/state
- Elasticsearch/OpenSearch only for search, not primary state
Metadata tables/entities:
- jobs
- url_state
- fetch_attempts
- page_versions
- link_graph
- content_dedup_index
Why split storage?
- Blob storage is cheap for large raw payloads
- DB is good for indexed metadata and scheduling
- Warehouse is good for analytics
9) Warehouse export design
The warehouse should receive clean, append-friendly tables.
Export options
- Batch export
- Periodically write Parquet/JSONL to object storage
- Load via COPY/LOAD jobs
- Streaming export
- Publish events to Kafka/PubSub
- Sink to warehouse via connector
- Hybrid
- Stream important metadata, batch raw crawl facts
Recommended warehouse tables
crawl_jobs
job_idsourcecrawl_typecreated_atschedule_typestatus
urls
url_idraw_urlnormalized_urldomainfirst_seen_atlast_seen_at
fetches
fetch_idurl_idjob_idfetched_athttp_statuscontent_typebytes_downloadedresponse_time_msetaglast_modifiedfinal_url
page_versions
page_version_idurl_idfetch_idcontent_hashsimhashtitletext_extractedcanonical_urlis_duplicatechanged_since_last
links
source_url_idtarget_urltarget_url_normalizedanchor_textreldiscovered_at
This structure lets analysts query crawl behavior, freshness, and content change over time.
10) Dedup strategy in practice
Use a layered approach.
Exact duplicate detection
- Compute hash on normalized content
- Example:
- strip whitespace
- remove boilerplate if desired
- lowercase if safe
- Use SHA-256 or BLAKE3
Near-duplicate detection
Use one of:
- SimHash
- MinHash + LSH
- document embeddings for semantic dedupe, if needed
Suggested policy
- If exact hash matches existing version for same URL: mark as unchanged
- If exact hash matches different URL: mark as duplicate content
- If near-duplicate threshold exceeded: mark as near-duplicate but keep sample if needed
Important: keep source URL lineage even when content is deduped.
11) Frontier design
The frontier is the heart of the crawler.
Requirements
- persistent
- distributed
- ordered by priority
- respects host constraints
- deduplicated
- resumable after failure
Frontier data model
You can think of it as:
- a priority queue of crawl tasks
- grouped by host/domain
- backed by durable storage
Suggested implementation pattern
- Task queue: Kafka/SQS/PubSub
- Host scheduler: DB/KV store
- Dedup index: Bloom filter + persistent state
- Worker lease system: tasks are leased, not permanently removed until acknowledged
Scheduling logic
When worker requests a task:
- select next eligible host
- check host crawl delay and concurrency
- pop next URL for that host
- lease it to worker with timeout
- on ack, mark done
- on timeout, requeue
12) Scaling approach
Horizontally scalable services
- many fetch workers
- many parser workers
- many export workers
Bottlenecks to design around
- frontier contention
- per-host rate limits
- database write amplification
- object storage upload throughput
- warehouse load latency
Techniques
- partition by domain hash
- batch writes
- async uploads
- caching robots.txt
- use worker autoscaling based on queue depth and fetch latency
- apply backpressure when warehouse/export lags
13) Reliability and correctness
Exactly-once is hard
At crawl scale, aim for at-least-once processing with idempotent writes.
Use idempotency keys like:
fetch_idurl_id + fetched_at_bucketcontent_hash
Failure handling
- worker crash: lease expires and task is retried
- partial upload: object storage upload is atomic from system perspective if using temp key + rename/copy
- duplicate task execution: dedupe at write time
Checkpointing
Keep:
- last processed offset in queues
- frontier state
- export watermark
- schedule watermark
14) Security and compliance
Include:
- robots.txt compliance
- allowlist/denylist
- crawl rate policies
- PII handling policy
- secrets management
- user-agent identification
- opt-out handling
- legal review for sensitive domains
If exporting content, consider:
- data retention limits
- redaction for sensitive fields
- access control in warehouse
15) Suggested technology stack
One possible stack:
Control plane
- API service: Go / Python / Java
- Scheduler: Airflow / Temporal / custom service
Queue/frontier
- Kafka for task stream
- Redis or RocksDB for dedupe/cache
- Postgres/Cassandra/DynamoDB for state
Workers
- Fetcher: Go, Rust, or Python with asyncio
- Browser rendering: Playwright/Chromium pool
- Parser: Python or JVM
Storage
- Raw content: S3/GCS
- Metadata: Postgres for small/medium, Cassandra/DynamoDB for large
- Search/index: OpenSearch if needed
Warehouse export
- Parquet files in object storage + Snowpipe/COPY/BigQuery load
- Or Kafka Connect / streaming sink
Orchestration/monitoring
- Kubernetes
- Prometheus + Grafana
- OpenTelemetry
- ELK / Loki
16) Minimal viable architecture
If you want to start simpler and evolve:
Phase 1
- Scheduler writes URLs to Kafka
- Workers fetch and store raw HTML in S3
- Postgres stores metadata
- Batch job writes parquet to warehouse
Phase 2
- Add URL normalization and exact dedupe
- Add robots and rate limiting
- Add conditional GET
- Add adaptive recrawl scheduling
Phase 3
- Add near-duplicate detection
- Add browser rendering
- Add host-aware frontier
- Add stronger distributed state
17) Example end-to-end workflow
- User submits
example.comwith weekly recrawl. - Scheduler seeds homepage and sitemap URLs.
- Frontier dedupes and prioritizes URLs.
- Worker fetches
/products. - Parser extracts 200 links.
- Normalization removes tracking parameters.
- 120 new URLs enter frontier.
- Page hash matches a previous version, so content is unchanged.
- Metadata records new fetch but no new page version.
- Nightly export writes crawl facts to Snowflake.
18) Practical schema sketch
url_state
url_idnormalized_urldomainfirst_seen_atlast_seen_atlast_crawled_atnext_crawl_atvisit_countfail_countlast_http_statuscontent_hashsimhashcanonical_url
crawl_task
task_idurl_idjob_idprioritylease_untilstatusattempt_count
fetch_result
fetch_idtask_idfinal_urlstatus_codeheaders_jsonraw_blob_uridownloaded_atduration_ms
extracted_link
source_url_idtarget_urltarget_normalizedanchor_textnofollowdiscovered_at
19) Key design choices to make early
Decide these up front:
- Scope
- whole web, a set of domains, or enterprise sites?
- Freshness
- how often should pages be revisited?
- Rendering
- do you need JS rendering?
- Dedup policy
- exact only, or near-duplicate too?
- Warehouse requirements
- batch or streaming?
- Compliance constraints
- robots, legal, PII, opt-outs
- Scale target
- 1M URLs/day vs 1B URLs/day changes everything
20) Recommended architecture summary
If I had to choose a solid default:
- Scheduler: Temporal or custom service
- Frontier: Kafka + durable KV store
- Dedup: URL normalization + Bloom filter + exact hash
- Fetch: Go workers with per-host token buckets
- Parsing: separate parser workers
- Storage: S3 for raw, Postgres/Cassandra for metadata
- Warehouse export: batch Parquet loads
- Recrawl: adaptive scheduling based on change history
- Monitoring: Prometheus/Grafana + structured logs
If you want, I can also provide any of these next:
- a diagram of the architecture,
- a database schema in SQL,
- a Kafka-based frontier design,
- a detailed scheduling algorithm,
- or a reference implementation in Python/Go.