Prompt

How do I dedupe URLs across a large crawl job?

Technology / Cloud Infrastructure1 observationLast seen Oct 4, 2026

Latest observation

Oct 4, 2026OpenAI APIWeb search: off

To dedupe URLs in a large crawl job, you want a pipeline that is:

  1. Canonicalization-first: normalize URLs so obvious duplicates become identical.
  2. Streaming-friendly: avoid holding the whole crawl in memory.
  3. Probabilistic + exact: use a fast approximate filter early, and exact storage for final decisions.
  4. Scope-aware: dedupe per crawl, per host, per domain, or globally depending on your goal.

1) Normalize URLs before deduping

A lot of duplicates are really formatting differences.

Common normalization steps:

  • Lowercase scheme and host
  • Remove default ports:
    • http:80, https:443
  • Normalize percent-encoding
  • Remove fragment identifiers (#section)
  • Normalize trailing slash rules consistently
  • Sort query parameters if order doesn’t matter
  • Drop tracking params like utm_*, gclid, etc. if appropriate
  • Resolve relative paths, . and ..
  • Convert punycode / IDN consistently

Example:

  • HTTPS://Example.com:443/a/b/../c/?b=2&a=1#top
  • https://example.com/a/c/?a=1&b=2

After normalization, these can be treated as the same URL.

2) Decide what “duplicate” means

There are several kinds:

  • Exact URL duplicates: same normalized URL string
  • Same page content: different URLs that serve identical content
  • Near duplicates: pages that are mostly the same
  • Crawl duplicates: already seen in this crawl job
  • Historical duplicates: seen in past crawl jobs

For URL dedupe during crawling, you usually want exact normalized URL dedupe first.

3) Use a two-stage dedupe design

For large crawls, a common pattern is:

Stage A: fast approximate filter

Use something like:

  • Bloom filter
  • Cuckoo filter
  • Xor filter

This answers: “Have I probably seen this URL before?”

Pros:

  • Very memory efficient
  • Very fast

Cons:

  • False positives possible
  • Not suitable as the only dedupe mechanism if missing URLs is bad

Stage B: exact set / persistent store

Use a durable store for exact dedupe decisions:

  • Redis set
  • RocksDB / LevelDB / LMDB
  • PostgreSQL / MySQL with unique index
  • Distributed KV store like DynamoDB / Bigtable / Cassandra
  • Sharded key-value store

This answers: “Have I definitely seen this URL before?”

A good architecture is:

  • Normalize URL
  • Hash it
  • Check Bloom filter
  • If “new”, check exact store and insert if absent
  • If Bloom says “seen”, optionally verify in exact store if false positives matter

4) Hash URLs instead of storing full strings in memory

Store a compact hash key rather than the full URL.

Recommended:

  • SHA-256 if you want collision safety
  • xxHash / MurmurHash for speed if collision risk is acceptable
  • Often store:
    • 64-bit hash for speed
    • plus the canonical URL in persistent store if needed

For very large-scale dedupe, a 128-bit or 256-bit hash is safer.

5) Partition/shard by hash

If the crawl is distributed, route URLs to workers by hash:

  • shard = hash(normalized_url) % N

This ensures:

  • All duplicates go to the same shard
  • Each worker maintains a local dedupe structure
  • Less cross-worker chatter

For massive scale, use:

  • Consistent hashing
  • Kafka topics partitioned by hash
  • Distributed KV store with key-based sharding

6) If you need “seen in this crawl job” only

Then you can use:

  • In-memory Bloom filter per worker for fast local dedupe
  • Shared persistent dedupe store for job-wide uniqueness

If job-wide uniqueness must be exact:

  • write every normalized URL hash to a centralized durable store with a unique constraint
  • or shard by hash and dedupe per shard

7) Handle URL scope correctly

Sometimes dedupe should be scoped differently:

  • Per-host: dedupe only within a site
  • Per-domain: dedupe within example.com
  • Global: dedupe across all crawls

If scoping matters, include the scope key in your dedupe key:

  • host + normalized_path
  • registrable_domain + normalized_url
  • crawl_job_id + normalized_url

8) Beware of query parameters

Query strings are a major source of duplicate-ish URLs.

You need a policy:

  • Keep all params
  • Sort params
  • Remove known tracking params
  • Whitelist only meaningful params

Example tracking params to remove:

  • utm_source
  • utm_medium
  • utm_campaign
  • gclid
  • fbclid

Be careful: some sites use query params for real content.

9) Consider content-based dedupe too

If you want to avoid crawling identical pages under different URLs:

  • Compute content hashes after fetch
  • Use SimHash / MinHash for near-duplicate detection
  • Keep a content fingerprint index

This is separate from URL dedupe.

10) Practical architecture example

A common production setup:

  1. Frontier emits candidate URLs
  2. Worker normalizes URL
  3. Worker hashes normalized URL
  4. Worker checks local Bloom filter
  5. If not seen locally:
    • check distributed exact store
    • if absent, insert and enqueue
  6. If seen, drop URL

For very large crawls:

  • Use a distributed dedupe store
  • Add a local cache/filter to reduce store load
  • Periodically compact / expire old entries if the crawl is time-bounded

11) If memory is tight

Use:

  • Bloom filter for first pass
  • Disk-backed key-value store for exact dedupe
  • Batch writes
  • Compression on stored keys

12) If you want a simple implementation

For a smaller crawl:

  • Normalize URLs
  • Store SHA-256 hashes in a database table with a unique index
  • Insert with “ignore if exists”
  • Use batching

Example schema:

CREATE TABLE seen_urls (
  url_hash BINARY(32) PRIMARY KEY,
  first_seen TIMESTAMP NOT NULL
);

Then:

  • compute normalized URL hash
  • insert
  • if insert succeeds, crawl it
  • if duplicate key, skip it

If you want, I can also give you:

  1. a Python implementation of URL normalization + Bloom filter dedupe,
  2. a distributed architecture for Kafka/Spark/Redis,
  3. or a database schema + insert pattern for exact dedupe.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.