Prompt

deduplicate with embeddings

Artificial Intelligence / AI Search1 observationLast seen Oct 10, 2026

Latest observation

Oct 10, 2026OpenAI APIWeb search: off

If you want to deduplicate records using embeddings, the basic idea is:

  1. Convert each item to an embedding vector
  2. Compare vectors by similarity
  3. Group items above a similarity threshold
  4. Keep one canonical record per group

Common approach

1) Generate embeddings

Use an embedding model for each text record, for example:

  • product names
  • company descriptions
  • support tickets
  • articles
  • addresses
from sentence_transformers import SentenceTransformer

model = SentenceTransformer("all-MiniLM-L6-v2")
embeddings = model.encode(texts, normalize_embeddings=True)

2) Find similar pairs

Use cosine similarity or nearest neighbors.

from sklearn.metrics.pairwise import cosine_similarity

sim_matrix = cosine_similarity(embeddings)

3) Cluster duplicates

Treat records as duplicates if similarity is above a threshold, e.g. 0.85.

threshold = 0.85
n = len(texts)
groups = []
visited = set()

for i in range(n):
    if i in visited:
        continue
    group = [i]
    visited.add(i)
    for j in range(i + 1, n):
        if sim_matrix[i][j] >= threshold:
            group.append(j)
            visited.add(j)
    groups.append(group)

4) Pick a representative

For each duplicate group, choose:

  • the shortest string
  • the most complete record
  • the earliest record
  • or use a rule-based canonicalization

Better for large datasets

For large datasets, don’t compute all pairs. Use:

  • FAISS
  • Annoy
  • ScaNN
  • approximate nearest neighbors

Example with FAISS:

import faiss
import numpy as np

emb = np.array(embeddings).astype("float32")
index = faiss.IndexFlatIP(emb.shape[1])  # for normalized embeddings
index.add(emb)

scores, ids = index.search(emb, k=5)  # top-5 nearest neighbors

Important caveats

Embeddings are semantic, not exact

They’re good for:

  • “iPhone 14 Pro Max 256GB”
  • “Apple iPhone 14 Pro Max 256 GB”

But not always for:

  • exact identifiers
  • part numbers
  • addresses with subtle differences
  • names where spelling matters

Combine with rules

Best practice is usually:

  • exact match first
  • then normalized text
  • then embedding similarity
  • then manual review for edge cases

Practical dedupe pipeline

A strong pipeline often looks like:

  1. Normalize text

    • lowercase
    • trim whitespace
    • remove punctuation
    • standardize abbreviations
  2. Block candidates

    • same first letter
    • same zip code
    • same domain
    • same category
  3. Embed candidate records

  4. Similarity search

  5. Threshold + clustering

  6. Canonical record selection


Example use case

For product deduplication:

  • "Nike Air Zoom Pegasus 40 Men's Running Shoes"
  • "Nike Pegasus 40 Air Zoom Running Shoe"

Embedding similarity may be high enough to group them, even though text differs.


If you want, I can give you:

  1. a Python script for deduplication with embeddings,
  2. a SQL + vector DB approach, or
  3. a production-grade pipeline design for large-scale dedupe.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.