Prompt
How do I dedupe URLs across a large crawl job?
Latest observation
To dedupe URLs in a large crawl job, you want a pipeline that is:
- Canonicalization-first: normalize URLs so obvious duplicates become identical.
- Streaming-friendly: avoid holding the whole crawl in memory.
- Probabilistic + exact: use a fast approximate filter early, and exact storage for final decisions.
- Scope-aware: dedupe per crawl, per host, per domain, or globally depending on your goal.
1) Normalize URLs before deduping
A lot of duplicates are really formatting differences.
Common normalization steps:
- Lowercase scheme and host
- Remove default ports:
http:80,https:443
- Normalize percent-encoding
- Remove fragment identifiers (
#section) - Normalize trailing slash rules consistently
- Sort query parameters if order doesn’t matter
- Drop tracking params like
utm_*,gclid, etc. if appropriate - Resolve relative paths,
.and.. - Convert punycode / IDN consistently
Example:
HTTPS://Example.com:443/a/b/../c/?b=2&a=1#tophttps://example.com/a/c/?a=1&b=2
After normalization, these can be treated as the same URL.
2) Decide what “duplicate” means
There are several kinds:
- Exact URL duplicates: same normalized URL string
- Same page content: different URLs that serve identical content
- Near duplicates: pages that are mostly the same
- Crawl duplicates: already seen in this crawl job
- Historical duplicates: seen in past crawl jobs
For URL dedupe during crawling, you usually want exact normalized URL dedupe first.
3) Use a two-stage dedupe design
For large crawls, a common pattern is:
Stage A: fast approximate filter
Use something like:
- Bloom filter
- Cuckoo filter
- Xor filter
This answers: “Have I probably seen this URL before?”
Pros:
- Very memory efficient
- Very fast
Cons:
- False positives possible
- Not suitable as the only dedupe mechanism if missing URLs is bad
Stage B: exact set / persistent store
Use a durable store for exact dedupe decisions:
- Redis set
- RocksDB / LevelDB / LMDB
- PostgreSQL / MySQL with unique index
- Distributed KV store like DynamoDB / Bigtable / Cassandra
- Sharded key-value store
This answers: “Have I definitely seen this URL before?”
A good architecture is:
- Normalize URL
- Hash it
- Check Bloom filter
- If “new”, check exact store and insert if absent
- If Bloom says “seen”, optionally verify in exact store if false positives matter
4) Hash URLs instead of storing full strings in memory
Store a compact hash key rather than the full URL.
Recommended:
- SHA-256 if you want collision safety
- xxHash / MurmurHash for speed if collision risk is acceptable
- Often store:
- 64-bit hash for speed
- plus the canonical URL in persistent store if needed
For very large-scale dedupe, a 128-bit or 256-bit hash is safer.
5) Partition/shard by hash
If the crawl is distributed, route URLs to workers by hash:
shard = hash(normalized_url) % N
This ensures:
- All duplicates go to the same shard
- Each worker maintains a local dedupe structure
- Less cross-worker chatter
For massive scale, use:
- Consistent hashing
- Kafka topics partitioned by hash
- Distributed KV store with key-based sharding
6) If you need “seen in this crawl job” only
Then you can use:
- In-memory Bloom filter per worker for fast local dedupe
- Shared persistent dedupe store for job-wide uniqueness
If job-wide uniqueness must be exact:
- write every normalized URL hash to a centralized durable store with a unique constraint
- or shard by hash and dedupe per shard
7) Handle URL scope correctly
Sometimes dedupe should be scoped differently:
- Per-host: dedupe only within a site
- Per-domain: dedupe within
example.com - Global: dedupe across all crawls
If scoping matters, include the scope key in your dedupe key:
host + normalized_pathregistrable_domain + normalized_urlcrawl_job_id + normalized_url
8) Beware of query parameters
Query strings are a major source of duplicate-ish URLs.
You need a policy:
- Keep all params
- Sort params
- Remove known tracking params
- Whitelist only meaningful params
Example tracking params to remove:
utm_sourceutm_mediumutm_campaigngclidfbclid
Be careful: some sites use query params for real content.
9) Consider content-based dedupe too
If you want to avoid crawling identical pages under different URLs:
- Compute content hashes after fetch
- Use SimHash / MinHash for near-duplicate detection
- Keep a content fingerprint index
This is separate from URL dedupe.
10) Practical architecture example
A common production setup:
- Frontier emits candidate URLs
- Worker normalizes URL
- Worker hashes normalized URL
- Worker checks local Bloom filter
- If not seen locally:
- check distributed exact store
- if absent, insert and enqueue
- If seen, drop URL
For very large crawls:
- Use a distributed dedupe store
- Add a local cache/filter to reduce store load
- Periodically compact / expire old entries if the crawl is time-bounded
11) If memory is tight
Use:
- Bloom filter for first pass
- Disk-backed key-value store for exact dedupe
- Batch writes
- Compression on stored keys
12) If you want a simple implementation
For a smaller crawl:
- Normalize URLs
- Store SHA-256 hashes in a database table with a unique index
- Insert with “ignore if exists”
- Use batching
Example schema:
CREATE TABLE seen_urls (
url_hash BINARY(32) PRIMARY KEY,
first_seen TIMESTAMP NOT NULL
);
Then:
- compute normalized URL hash
- insert
- if insert succeeds, crawl it
- if duplicate key, skip it
If you want, I can also give you:
- a Python implementation of URL normalization + Bloom filter dedupe,
- a distributed architecture for Kafka/Spark/Redis,
- or a database schema + insert pattern for exact dedupe.