Prompt

How do I use embeddings for deduplication of near-duplicate records?

Artificial Intelligence / AI Search1 observationLast seen Oct 10, 2026

Latest observation

Oct 10, 2026OpenAI APIWeb search: off

A good way to use embeddings for deduplication is to treat each record as a point in vector space, then find records that are very close together and verify they’re truly duplicates.

High-level approach

  1. Embed each record

    • Convert the text or structured fields into an embedding vector.
    • For structured records, you can combine fields into a canonical text form, e.g.:
      • "Acme Corp | 123 Main St | New York | 10001"
    • For better results, embed important fields separately and concatenate or compare them with weighting.
  2. Choose a similarity measure

    • Usually cosine similarity for embeddings.
    • Near-duplicates will have high cosine similarity.
  3. Find candidates efficiently

    • Don’t compare every record to every other record if you have many records.
    • Use:
      • Approximate nearest neighbors (ANN) via FAISS, Milvus, Pinecone, pgvector, etc.
      • A brute-force scan if the dataset is small.
  4. Set a threshold

    • Decide a similarity cutoff above which records are considered likely duplicates.
    • Example: cosine_similarity > 0.90 might be a starting point, but you must tune it on labeled examples.
  5. Cluster or pairwise merge

    • If A is similar to B and B to C, use clustering or connected components to form duplicate groups.
    • Then either:
      • merge automatically, or
      • send to human review if confidence is uncertain.
  6. Add rule-based checks

    • Embeddings are great for semantic similarity, but you often want additional signals:
      • exact email/phone match
      • normalized address match
      • edit distance on names
      • same product ID / invoice number
    • Combine embedding similarity with deterministic rules for better precision.

Typical workflow

For a batch dedup job

  1. Normalize records:

    • lowercase
    • trim whitespace
    • standardize punctuation
    • normalize dates/phones/addresses if applicable
  2. Create embeddings for each record.

  3. Build an ANN index.

  4. For each record, retrieve top-k nearest neighbors.

  5. Filter candidates using:

    • similarity threshold
    • business rules
    • maybe a second-stage classifier
  6. Merge duplicates into groups.


Example

Suppose you have:

  • “Apple iPhone 15 Pro Max 256GB Black”
  • “Apple iPhone 15 ProMax 256 GB - Black”
  • “Samsung Galaxy S24 Ultra”

The first two should likely be close in embedding space, while the third should be far away.


Practical tips

1. Use embeddings for candidate generation, not final truth

Embeddings are good for finding likely matches, but not perfect. A second-stage check improves quality.

2. Calibrate thresholds with labeled data

Create a small labeled set:

  • duplicate pairs
  • non-duplicate pairs

Then inspect similarity distributions and pick a threshold that balances precision and recall.

3. Beware of short records

Very short strings like names or codes can be ambiguous. For these, embeddings alone may not be enough.

4. Consider field-aware embeddings

Instead of embedding the whole record as one string, you can:

  • embed name, address, and description separately
  • compare each field differently
  • combine scores

5. Re-embed when the data schema changes

If you change field formatting or add new fields, rebuild embeddings and re-tune thresholds.


Simple implementation pattern

Pseudo-code:

for record in records:
    text = normalize(record)
    vector = embed(text)
    index.add(record_id, vector)

for record in records:
    neighbors = index.search(embed(normalize(record)), top_k=20)
    for n in neighbors:
        if cosine_similarity(record.vector, n.vector) > threshold:
            mark_as_possible_duplicate(record, n)

When embeddings work well

  • product catalogs
  • customer support tickets
  • business names
  • document titles/descriptions
  • messy free text with minor wording differences

When they work less well

  • exact identifiers matter more than semantics
  • tiny strings
  • records differing mainly in numeric fields
  • cases requiring strict legal/financial accuracy

If you want, I can also show:

  1. a Python example with FAISS or pgvector, or
  2. a more robust dedup pipeline combining embeddings + rules + clustering.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.