Prompt
I'm building a deduplication workflow and want to use embeddings to catch near-duplicates
Latest observation
Great use case. Embeddings work well for near-duplicate detection when you want to catch:
- exact duplicates with minor edits
- paraphrases
- formatting changes
- reordered but semantically similar content
A practical dedup workflow usually looks like this:
1) Normalize the text first
Before embedding, apply lightweight cleanup:
- lowercase
- strip extra whitespace
- remove boilerplate if possible
- normalize punctuation
- optionally remove HTML/markup
This helps reduce noise and improves consistency.
2) Generate embeddings
Embed each document, sentence, or chunk depending on your granularity.
- Whole document embeddings: good if records are short or mostly uniform
- Chunk embeddings: better for long documents where duplicates may appear in sections
3) Find nearest neighbors
Use cosine similarity or dot product on normalized vectors.
Common approaches:
- All-pairs search for small datasets
- ANN index for large datasets:
- FAISS
- hnswlib
- vector databases
4) Set a similarity threshold
You’ll need to tune this empirically.
Typical starting point:
>= 0.95: likely duplicates0.85–0.95: near-duplicates / review bucket< 0.85: probably different
Exact thresholds vary a lot by model and content type.
5) Add a second-stage check
Embeddings are great for candidate generation, but for high precision, follow with:
- token-based similarity
- edit distance
- MinHash / shingling
- exact field comparisons
This reduces false positives.
6) Cluster or link duplicates
Depending on your needs:
- pairwise dedupe: mark likely duplicate pairs
- entity clustering: group multiple records into duplicate clusters
- canonical selection: choose one record as the survivor
7) Human review for borderline cases
For records near the threshold, route them to manual review.
Recommended architecture
A strong pattern is:
- Normalize
- Embed
- ANN retrieve top-k candidates
- Score with cosine similarity
- Optional lexical re-rank
- Apply threshold / cluster
- Review borderline pairs
Example
If two records are:
- “Apple iPhone 15 Pro Max 256GB, natural titanium”
- “Apple iPhone 15 Pro Max, 256 GB, Natural Titanium”
They’ll likely embed very close even though punctuation and spacing differ.
Important caveats
Embeddings can struggle with:
- short strings like IDs, SKUs, phone numbers
- highly structured records where small changes matter
- content where exact wording is critical
In those cases, combine embeddings with exact-match or rule-based logic.
If you want, I can help you design:
- a dedupe pipeline,
- thresholding and evaluation strategy, or
- a Python example using embeddings + FAISS/hnswlib.