Prompt
what should I use for deduplication with vector similarity?
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
For deduplication with vector similarity, a good default is:
Use embeddings + nearest-neighbor search + a similarity threshold
Typical stack
- Embeddings: convert each item to a vector
- ANN index: FAISS, hnswlib, ScaNN, or a vector DB
- Metric: usually cosine similarity
- Decision rule: if the nearest neighbor is above a threshold, treat as duplicate
What to choose
If you want simple and fast
- FAISS if you’re doing this in Python and want control/performance
- hnswlib if you want very good approximate search and easy setup
If you’re already using a database/vector platform
- Pinecone, Weaviate, Milvus, Qdrant, Elasticsearch vector search, PostgreSQL + pgvector
Recommended approach
- Generate embeddings for all items
- Normalize vectors if using cosine similarity
- For each new item:
- search nearest neighbors
- if top match similarity > threshold, mark as duplicate
- Optionally also require:
- same source/domain
- high textual overlap
- exact metadata match on key fields
Threshold guidance
This depends heavily on the embedding model and data type, but roughly:
- 0.95+: near-exact duplicates
- 0.85–0.95: probably duplicates, review if risky
- below 0.85: often similar but not duplicates
You should calibrate on a labeled sample.
Important caveats
- Vector similarity is good for semantic near-duplicates, not exact dedup only.
- For exact duplicates, use:
- hashes
- normalization
- string similarity
- Best results usually come from a hybrid approach:
- exact match first
- vector similarity second
- optional rule-based filters third
My default recommendation
If you’re asking “what should I use” and want a practical answer:
- Embeddings + cosine similarity + HNSW/FAISS
- Add a threshold tuned on sample data
- Use exact hashing for strict duplicates
If you tell me:
- what kind of data you’re deduplicating
- scale
- batch vs real-time
- language/domain
…I can recommend a specific threshold and tool.