Prompt
deduplicate with embeddings
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
If you want to deduplicate records using embeddings, the basic idea is:
- Convert each item to an embedding vector
- Compare vectors by similarity
- Group items above a similarity threshold
- Keep one canonical record per group
Common approach
1) Generate embeddings
Use an embedding model for each text record, for example:
- product names
- company descriptions
- support tickets
- articles
- addresses
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("all-MiniLM-L6-v2")
embeddings = model.encode(texts, normalize_embeddings=True)
2) Find similar pairs
Use cosine similarity or nearest neighbors.
from sklearn.metrics.pairwise import cosine_similarity
sim_matrix = cosine_similarity(embeddings)
3) Cluster duplicates
Treat records as duplicates if similarity is above a threshold, e.g. 0.85.
threshold = 0.85
n = len(texts)
groups = []
visited = set()
for i in range(n):
if i in visited:
continue
group = [i]
visited.add(i)
for j in range(i + 1, n):
if sim_matrix[i][j] >= threshold:
group.append(j)
visited.add(j)
groups.append(group)
4) Pick a representative
For each duplicate group, choose:
- the shortest string
- the most complete record
- the earliest record
- or use a rule-based canonicalization
Better for large datasets
For large datasets, don’t compute all pairs. Use:
- FAISS
- Annoy
- ScaNN
- approximate nearest neighbors
Example with FAISS:
import faiss
import numpy as np
emb = np.array(embeddings).astype("float32")
index = faiss.IndexFlatIP(emb.shape[1]) # for normalized embeddings
index.add(emb)
scores, ids = index.search(emb, k=5) # top-5 nearest neighbors
Important caveats
Embeddings are semantic, not exact
They’re good for:
- “iPhone 14 Pro Max 256GB”
- “Apple iPhone 14 Pro Max 256 GB”
But not always for:
- exact identifiers
- part numbers
- addresses with subtle differences
- names where spelling matters
Combine with rules
Best practice is usually:
- exact match first
- then normalized text
- then embedding similarity
- then manual review for edge cases
Practical dedupe pipeline
A strong pipeline often looks like:
-
Normalize text
- lowercase
- trim whitespace
- remove punctuation
- standardize abbreviations
-
Block candidates
- same first letter
- same zip code
- same domain
- same category
-
Embed candidate records
-
Similarity search
-
Threshold + clustering
-
Canonical record selection
Example use case
For product deduplication:
"Nike Air Zoom Pegasus 40 Men's Running Shoes""Nike Pegasus 40 Air Zoom Running Shoe"
Embedding similarity may be high enough to group them, even though text differs.
If you want, I can give you:
- a Python script for deduplication with embeddings,
- a SQL + vector DB approach, or
- a production-grade pipeline design for large-scale dedupe.