Prompt
How do I set up a site ingestion platform for crawling public pages with deduplication and canonicalization?
Latest observation
A solid site ingestion platform for crawling public web pages usually has 5 layers:
- Discovery / crawl frontier
- Fetch / render
- Canonicalization
- Deduplication
- Storage / indexing / monitoring
Below is a practical setup that works well for large-scale crawling.
1) Define the crawl scope and rules
Before building anything, decide:
- Allowed domains / subdomains
- Respect robots.txt and crawl-delay where applicable
- Rate limits per host
- Content types to ingest: HTML only, or also PDFs, images, feeds, etc.
- Freshness policy: full recrawl, incremental updates, or both
- Language / region filters
- Auth boundaries: public only, no login walls
For public web crawling, you should also:
- Identify your crawler with a clear User-Agent
- Provide a contact email / website
- Respect robots exclusions
- Avoid aggressive concurrency on single hosts
2) Build the crawl frontier
The frontier is the queue of URLs to fetch.
Core frontier features
- URL normalization before enqueueing
- Per-host politeness
- Priority scheduling
- Duplicate URL suppression
- Seen URL tracking
- Freshness-aware re-crawl scheduling
Typical components
- A queue or topic system:
- Redis / Kafka / RabbitMQ / SQS
- A URL fingerprint store:
- Bloom filter for fast approximate checks
- Persistent key-value store for exact “seen” status
- A scheduler:
- Assigns URLs to workers based on host, priority, and crawl rules
Good frontier policy
- Prefer breadth-first for discovery
- Revisit important pages more frequently
- Throttle per host and per IP
- Group URLs by host to reduce connection overhead
3) Fetch and render pages
Fetching
Use an HTTP client that supports:
- Redirect handling
- Compression
- Cookies when needed
- Timeouts and retries
- Conditional requests:
If-Modified-SinceIf-None-Match
Rendering
Some pages require JavaScript. Use a dual pipeline:
- Fast path: raw HTTP fetch for normal pages
- Rendered path: headless browser for JS-heavy pages
Typical tools:
- HTTP: Go, Python requests/httpx, Node fetch, Rust reqwest
- Rendering: Playwright, Puppeteer, Selenium
Store raw response metadata
Keep:
- final URL after redirects
- status code
- response headers
- fetch timestamp
- content type
- robots decision
- fetch duration
- render flags
- hash of raw body
4) Canonicalization: normalize URLs and content
Canonicalization is key to avoiding duplicates and consistent storage.
A. URL canonicalization
Normalize URLs before enqueueing and before deduping.
Common rules
- Lowercase scheme and host
- Remove default ports:
:80for HTTP:443for HTTPS
- Normalize trailing slashes consistently
- Remove URL fragments (
#...) - Sort query parameters when safe
- Remove known tracking parameters:
utm_*gclidfbclidref
- Decode percent-encoding where safe
- Convert punycode/IDN consistently
- Resolve relative links against base URL
- Normalize path dot segments:
/a/./b/../c→/a/c
Be careful with query params
Do not blindly sort/remove all query params. Some are semantically important. Use:
- per-domain rules, or
- allowlist/denylist policies
Canonical URL source
If HTML includes:
<link rel="canonical" href="...">
Use it as a signal, not absolute truth.
A good approach:
- Parse canonical link
- Normalize it
- Trust it only if:
- same host or approved cross-host relationship
- page content strongly matches
- Store both:
- fetched URL
- declared canonical URL
- resolved canonical URL
B. Content canonicalization
To dedupe content, normalize the page body after extraction.
For HTML pages
Extract the main content if you want document-level dedupe:
- Remove scripts, styles, nav, boilerplate
- Normalize whitespace
- Normalize entities
- Lowercase only if appropriate for text matching
- Strip hidden/duplicate elements
Tools:
- Readability
- Boilerpipe
- trafilatura
- BeautifulSoup/lxml-based extraction
For general content dedupe
Create one or more hashes:
- Exact hash of raw body
- Normalized text hash
- Near-duplicate fingerprint
- SimHash
- MinHash / LSH
- shingling with w-shingles
A practical approach:
raw_hashfor exact byte-level duplicatesnormalized_text_hashfor same content with formatting differencessimhashfor near-duplicates
5) Deduplication strategy
Use dedupe at multiple levels.
A. URL-level dedupe
Prevent fetching the same normalized URL twice.
Store:
- normalized URL hash
- last seen timestamp
- fetch status
- canonical target if known
B. Document-level exact dedupe
After fetch:
- compute
raw_hashand/ornormalized_hash - if hash already exists, mark as duplicate and skip storage of full content if desired
C. Near-duplicate dedupe
Pages may differ slightly:
- ad changes
- timestamps
- minor layout changes
Use:
- SimHash for efficient similarity comparison
- MinHash/LSH for clustering similar documents
Typical threshold:
- SimHash Hamming distance under a small threshold means near-duplicate
- Tune based on observed content types
Dedupe flow
- Normalize URL
- Check frontier seen set
- Fetch page
- Extract canonical URL
- Normalize content
- Compute hashes/fingerprints
- Compare against dedupe index
- Store or mark duplicate
6) Storage design
A robust storage model usually separates:
Raw fetch store
Keep original response for replay/debugging:
- object storage: S3, GCS, Azure Blob, MinIO
Store:
- HTML/body
- headers
- metadata JSON
- screenshot or rendered DOM if needed
Metadata database
Use Postgres, MySQL, or a document store.
Tables/collections:
pagesfetchesurlscanonical_mappingscontent_fingerprintscrawl_jobs
Search/index store
If you need retrieval/search:
- Elasticsearch / OpenSearch / Solr / Vespa / BigQuery / ClickHouse
- store extracted text, titles, entities, links, hashes
7) A practical data model
urls
idnormalized_urlurl_hashhostschemefirst_seen_atlast_seen_atcrawl_status
fetches
idurl_idfetched_atstatus_codecontent_typefinal_urletaglast_modifiedbody_object_keyraw_hashrenderedbool
documents
idurl_idfetch_idcanonical_urltitleextracted_textnormalized_text_hashsimhashlanguagecontent_length
links
source_url_idtarget_urltarget_normalized_urlanchor_textnofollowdiscovered_at
8) Crawl scheduling and recrawl policy
Basic scheduling
- New URLs go to frontier
- Frequent recrawl for high-value pages
- Back off on repeated errors
- Respect
304 Not Modified
Prioritization signals
- URL path depth
- site importance
- change frequency
- inbound link count
- sitemap presence
- content type
Common recrawl patterns
- News: minutes to hours
- Blogs/docs: daily to weekly
- Static reference pages: weekly to monthly
9) Handling edge cases
Redirect chains
- Record all hops
- Deduplicate on final normalized destination
- Keep source-to-target mapping
Canonical conflicts
If page A says canonical B, but page B says canonical C:
- prefer same-host canonical consistency
- rely on content similarity
- maintain canonical graph and resolve cycles
Session IDs and traps
Avoid infinite URL spaces:
- remove session params
- detect calendar/trap patterns
- limit path depth and parameter combinations
- cap pages per host/day
Duplicate templates
Many sites have boilerplate pages:
- category pages
- tag pages
- printer-friendly views
- AMP pages
Decide whether to:
- ingest all variants
- collapse to one canonical representation
- keep variant type as metadata
10) Suggested pipeline architecture
Simple scalable setup
- Crawler workers: fetch URLs
- Scheduler/frontier service: manages queues and politeness
- Parser/canonicalizer service: extracts links, content, canonical URL
- Dedup service: hashes and similarity checks
- Storage service: raw blobs + metadata DB
- Indexer: pushes text into search system
- Monitoring: metrics and logs
Message flow
- Seed URLs enter frontier
- Scheduler assigns URL to worker
- Worker fetches page
- Parser extracts links and canonical info
- New links normalized and re-enqueued
- Dedup service checks hashes
- Document stored or marked duplicate
11) Recommended implementation choices
If you want something fast to build
- Python for orchestration/parsing
- Playwright for rendering
- Redis for frontier/locking
- Postgres for metadata
- S3/MinIO for raw content
- OpenSearch for search
- SimHash for near-duplicate detection
If you want higher throughput
- Go or Rust for fetchers and URL normalization
- Kafka for work distribution
- Postgres/ClickHouse for metadata analytics
- S3 for storage
- LSH/SimHash service for dedupe
12) Minimal canonicalization rules to start with
Start with a conservative set:
- lower-case scheme + host
- remove fragment
- remove default ports
- normalize path
- remove
utm_*,gclid,fbclid - resolve relative URLs
- respect
<link rel="canonical"> - compute exact and normalized text hashes
Then add domain-specific rules as you observe patterns.
13) Testing and validation
Test against:
- URLs with redirects
- duplicate pages with different params
- pages with and without canonical tags
- dynamic JS pages
- pages with malformed HTML
- infinite parameter spaces
Measure:
- duplicate fetch rate
- duplicate content rate
- canonical agreement rate
- 304 hit rate
- crawl politeness violations
- extraction failure rate
14) A good first MVP
If you’re starting from scratch, build this first:
- Seed URL list
- URL normalization
- Robots check
- Polite HTTP fetch
- Raw body hashing
- HTML text extraction
- Canonical tag extraction
- Exact duplicate detection
- Link extraction and frontier enqueue
- Store everything in Postgres + object storage
Then add:
- rendered crawling
- SimHash near-duplicate detection
- advanced scheduling
- search indexing
If you want, I can also give you:
- a reference architecture diagram
- a Postgres schema
- a Python/Go implementation outline
- or a dedup/canonicalization algorithm in pseudocode.