Prompt

I’m building semantic clustering for feedback data and need advice on embeddings

Artificial Intelligence / AI Search1 observationLast seen Oct 10, 2026

Latest observation

Oct 10, 2026OpenAI APIWeb search: off

For semantic clustering of feedback data, the best embedding choice depends on your data shape, scale, and clustering goal. Here’s a practical framework:

1) Start with your use case

If feedback is short and messy

Examples: app reviews, survey comments, support tickets, NPS verbatims
Use sentence/document embeddings, not word embeddings.

Good options:

  • OpenAI text-embedding-3-small / large
  • Sentence-BERT / all-MiniLM-L6-v2
  • E5 / BGE family for retrieval-style semantics
  • Instructor models if you want task-specific prompts

If you need high-quality semantic grouping

Prioritize:

  • strong sentence-level meaning
  • robustness to paraphrases
  • multilingual support if needed
  • low sensitivity to formatting and punctuation

2) Recommended embedding strategy

Best general-purpose approach

  1. Embed each feedback item
  2. Reduce dimensionality only if needed
  3. Cluster with a density-based or hierarchical method
  4. Label clusters using representative examples

A strong default pipeline:

  • Embedding model: text-embedding-3-small or text-embedding-3-large
  • Optional dimensionality reduction: UMAP
  • Clustering: HDBSCAN or Agglomerative Clustering
  • Inspection: top exemplars + centroid nearest neighbors

3) Which embedding model to choose

OpenAI embeddings

Pros:

  • strong semantic quality
  • easy API usage
  • good for clustering, search, deduplication

Use:

  • text-embedding-3-small for cost-efficient large-scale clustering
  • text-embedding-3-large for better separation if budget allows

Sentence-BERT / open-source

Pros:

  • self-hosting
  • no per-call cost
  • good control over deployment

Popular choices:

  • all-MiniLM-L6-v2 for speed
  • bge-base/en or bge-large
  • e5-base / e5-large

These are great if you want local deployment or need to cluster millions of records cheaply.

4) Clustering methods that work well

HDBSCAN

Best when:

  • cluster count is unknown
  • lots of noise/outliers
  • uneven cluster sizes

Agglomerative clustering

Best when:

  • you want interpretable hierarchy
  • you have some sense of granularity
  • you can choose a similarity threshold

K-means

Best when:

  • you know the approximate number of clusters
  • you want simplicity and speed

For feedback data, HDBSCAN + embeddings is often the most robust starting point.

5) Preprocessing tips

Before embedding:

  • keep the original text mostly intact
  • remove obvious boilerplate if repeated everywhere
  • normalize whitespace
  • consider preserving emojis if sentiment matters
  • don’t over-clean punctuation/casing unless data is extremely noisy

Avoid heavy stemming/lemmatization for embeddings; modern models usually handle raw text better.

6) Handling long feedback

If comments are long:

  • split into chunks by sentence or paragraph
  • embed chunks separately
  • average chunk embeddings or store them individually
  • for clustering, one comment = one embedding usually works unless comments are very long and multi-topic

7) Measuring whether embeddings are good

Use:

  • manual cluster review
  • silhouette score cautiously
  • cluster purity on labeled subsets
  • neighbor sanity checks: do nearest neighbors look semantically similar?

For feedback, human inspection is often more useful than a single numeric score.

8) Practical recommendation

If you want a strong default:

  • Embeddings: text-embedding-3-small
  • Clustering: UMAP + HDBSCAN
  • Cluster interpretation: nearest examples + optional LLM-generated labels

If you need maximum quality:

  • Embeddings: text-embedding-3-large
  • Clustering: HDBSCAN or hierarchical
  • Labeling: LLM-assisted summaries of each cluster

9) A common workflow

  1. Clean feedback lightly
  2. Generate embeddings
  3. Reduce with UMAP if dataset is large/high-dimensional
  4. Cluster with HDBSCAN
  5. Inspect cluster exemplars
  6. Merge/split manually if needed
  7. Iterate with thresholds and parameters

10) If you want, I can help with one of these next:

  • recommend the best embedding model for your specific data
  • give you a Python example pipeline
  • suggest clustering parameters
  • compare OpenAI vs open-source embeddings
  • help design a labeling taxonomy for feedback clusters

If you share:

  • sample feedback text
  • language(s)
  • data size
  • whether you need self-hosting
  • desired number of clusters or granularity

…I can give a much more specific recommendation.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.