Prompt

How do I use embeddings for deduping near-duplicate records?

Artificial Intelligence / AI Search1 observationLast seen Oct 10, 2026

Latest observation

Oct 10, 2026OpenAI APIWeb search: off

A common pattern is:

  1. Turn each record into text
    Example: concatenate fields into a normalized string, like
    "Acme Corp | 123 Main St | Springfield | acme@example.com"

  2. Create embeddings for each record
    Use an embedding model to map each text record to a vector.

  3. Compare vectors by cosine similarity
    Near-duplicates should have high similarity.

  4. Choose a threshold
    If similarity is above a threshold, treat records as potential duplicates.

  5. Cluster or pair records
    Instead of comparing every record to every other record, use approximate nearest neighbors or clustering to find candidate duplicates efficiently.

Practical workflow

1) Normalize first

Embeddings work much better if you clean obvious noise:

  • lowercase
  • trim whitespace
  • standardize abbreviations
  • remove punctuation if appropriate
  • normalize phone numbers, addresses, dates
  • canonicalize known values

2) Build a “record text”

If you have structured data, don’t just dump raw JSON. Make a consistent text template:

name: Acme Corp
address: 123 Main St, Springfield, IL
email: support@acme.com
phone: 555-123-4567

This helps the model keep field meaning.

3) Embed each record

Store:

  • record ID
  • embedding vector
  • maybe some metadata for rule-based checks

4) Retrieve candidates

For a new record:

  • compute its embedding
  • find top-k nearest neighbors in your vector index
  • compare only those candidates more carefully

5) Combine embeddings with rules

Embeddings are good for fuzzy similarity, but deduping is usually best as a hybrid:

  • exact match rules for email/phone/customer ID
  • embedding similarity for messy names/addresses/descriptions
  • secondary deterministic checks before merging

Example logic

If similarity > 0.92:

  • auto-merge

If 0.80–0.92:

  • flag for review

If < 0.80:

  • treat as distinct

These thresholds are not universal; you must tune them on labeled examples.


Important caveats

Embeddings are not perfect for:

  • exact identity matching
  • numeric fields that must match exactly
  • records where tiny differences matter a lot
  • very short strings like codes or IDs

Better to use embeddings when:

  • names vary slightly
  • addresses are messy
  • descriptions are paraphrased
  • records contain semi-structured text

For best results:

  • use a blocking strategy first, such as same country, same first letter, or same zip prefix
  • then use embeddings on the reduced candidate set
  • then run a final deterministic or model-based match

Minimal Python sketch

from sklearn.metrics.pairwise import cosine_similarity
import numpy as np

# embeddings: list of vectors for existing records
# new_emb: vector for new record

sims = cosine_similarity([new_emb], embeddings)[0]
top_idx = np.argsort(sims)[::-1][:5]

for idx in top_idx:
    print(idx, sims[idx])

For scale, use a vector database or ANN library like:

  • FAISS
  • Annoy
  • HNSWlib
  • pgvector
  • Pinecone / Weaviate / Milvus

Best practice summary

  • Normalize data first
  • Embed a well-structured textual representation
  • Use nearest-neighbor search to get candidates
  • Tune similarity thresholds on labeled examples
  • Combine embeddings with exact rules and business logic

If you want, I can also show:

  1. a full dedupe pipeline design,
  2. a Python example using FAISS, or
  3. how to choose thresholds with labeled pairs.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.