Prompt
I’m building semantic clustering for feedback data and need advice on embeddings
Latest observation
For semantic clustering of feedback data, the best embedding choice depends on your data shape, scale, and clustering goal. Here’s a practical framework:
1) Start with your use case
If feedback is short and messy
Examples: app reviews, survey comments, support tickets, NPS verbatims
Use sentence/document embeddings, not word embeddings.
Good options:
- OpenAI text-embedding-3-small / large
- Sentence-BERT / all-MiniLM-L6-v2
- E5 / BGE family for retrieval-style semantics
- Instructor models if you want task-specific prompts
If you need high-quality semantic grouping
Prioritize:
- strong sentence-level meaning
- robustness to paraphrases
- multilingual support if needed
- low sensitivity to formatting and punctuation
2) Recommended embedding strategy
Best general-purpose approach
- Embed each feedback item
- Reduce dimensionality only if needed
- Cluster with a density-based or hierarchical method
- Label clusters using representative examples
A strong default pipeline:
- Embedding model:
text-embedding-3-smallortext-embedding-3-large - Optional dimensionality reduction: UMAP
- Clustering: HDBSCAN or Agglomerative Clustering
- Inspection: top exemplars + centroid nearest neighbors
3) Which embedding model to choose
OpenAI embeddings
Pros:
- strong semantic quality
- easy API usage
- good for clustering, search, deduplication
Use:
text-embedding-3-smallfor cost-efficient large-scale clusteringtext-embedding-3-largefor better separation if budget allows
Sentence-BERT / open-source
Pros:
- self-hosting
- no per-call cost
- good control over deployment
Popular choices:
all-MiniLM-L6-v2for speedbge-base/enorbge-largee5-base/e5-large
These are great if you want local deployment or need to cluster millions of records cheaply.
4) Clustering methods that work well
HDBSCAN
Best when:
- cluster count is unknown
- lots of noise/outliers
- uneven cluster sizes
Agglomerative clustering
Best when:
- you want interpretable hierarchy
- you have some sense of granularity
- you can choose a similarity threshold
K-means
Best when:
- you know the approximate number of clusters
- you want simplicity and speed
For feedback data, HDBSCAN + embeddings is often the most robust starting point.
5) Preprocessing tips
Before embedding:
- keep the original text mostly intact
- remove obvious boilerplate if repeated everywhere
- normalize whitespace
- consider preserving emojis if sentiment matters
- don’t over-clean punctuation/casing unless data is extremely noisy
Avoid heavy stemming/lemmatization for embeddings; modern models usually handle raw text better.
6) Handling long feedback
If comments are long:
- split into chunks by sentence or paragraph
- embed chunks separately
- average chunk embeddings or store them individually
- for clustering, one comment = one embedding usually works unless comments are very long and multi-topic
7) Measuring whether embeddings are good
Use:
- manual cluster review
- silhouette score cautiously
- cluster purity on labeled subsets
- neighbor sanity checks: do nearest neighbors look semantically similar?
For feedback, human inspection is often more useful than a single numeric score.
8) Practical recommendation
If you want a strong default:
- Embeddings:
text-embedding-3-small - Clustering: UMAP + HDBSCAN
- Cluster interpretation: nearest examples + optional LLM-generated labels
If you need maximum quality:
- Embeddings:
text-embedding-3-large - Clustering: HDBSCAN or hierarchical
- Labeling: LLM-assisted summaries of each cluster
9) A common workflow
- Clean feedback lightly
- Generate embeddings
- Reduce with UMAP if dataset is large/high-dimensional
- Cluster with HDBSCAN
- Inspect cluster exemplars
- Merge/split manually if needed
- Iterate with thresholds and parameters
10) If you want, I can help with one of these next:
- recommend the best embedding model for your specific data
- give you a Python example pipeline
- suggest clustering parameters
- compare OpenAI vs open-source embeddings
- help design a labeling taxonomy for feedback clusters
If you share:
- sample feedback text
- language(s)
- data size
- whether you need self-hosting
- desired number of clusters or granularity
…I can give a much more specific recommendation.