Prompt
How do I use embeddings for clustering similar tickets or articles?
Latest observation
Embeddings are a good way to turn tickets or articles into numeric vectors so you can cluster by semantic similarity instead of just keyword overlap.
High-level approach
-
Prepare the text
- Combine useful fields into one string.
- Examples:
- Support ticket:
subject + description + resolution - Article:
title + summary + body
- Support ticket:
- Clean obvious noise, but don’t over-process.
-
Generate embeddings
- Use an embedding model to convert each ticket/article into a vector.
- Similar texts should have nearby vectors.
-
Optionally reduce dimensionality
- Embeddings are often 384–3072 dimensions.
- For clustering, you may use them directly or reduce with PCA/UMAP.
-
Cluster the vectors
- Common choices:
- K-Means if you know roughly how many clusters you want
- Hierarchical clustering if you want nested groups
- DBSCAN/HDBSCAN if you want automatic cluster count and outlier handling
- Common choices:
-
Inspect cluster quality
- Look at sample items from each cluster.
- Use cluster labels/centroids to summarize themes.
Typical pipeline
1) Create text input
Example for tickets:
Title: Password reset not working
Body: I cannot reset my password. The reset email never arrives.
Category: Login
2) Embed the text
Using an embedding API or library, get a vector for each item.
3) Cluster
For example, with Python + scikit-learn:
from sklearn.cluster import KMeans
X = embeddings # shape: (n_items, embedding_dim)
kmeans = KMeans(n_clusters=10, random_state=42)
labels = kmeans.fit_predict(X)
4) Review clusters
For each cluster:
- print top examples
- find representative items closest to the centroid
Choosing a clustering method
K-Means
Best if:
- you have a rough sense of cluster count
- clusters are fairly spherical in embedding space
Pros:
- fast, simple Cons:
- requires choosing
k
HDBSCAN / DBSCAN
Best if:
- you don’t know the number of clusters
- you expect noise/outliers
Pros:
- finds natural groupings
- can mark outliers Cons:
- more sensitive to parameters
- can struggle if clusters have very different densities
Hierarchical clustering
Best if:
- you want a tree of topics or subtopics
Pros:
- interpretable Cons:
- can be slower on large datasets
Important tips
Use cosine similarity
For text embeddings, cosine similarity is often better than Euclidean distance.
If you use K-Means, many workflows still work well, especially if embeddings are normalized first.
Normalize embeddings
Often helpful:
from sklearn.preprocessing import normalize
X = normalize(embeddings)
Combine multiple fields carefully
For tickets, good inputs might be:
- subject
- description
- tags
- resolution notes
But avoid including fields that leak labels or add noise unless they help.
Cluster at the right granularity
- Too coarse: one cluster contains many unrelated issues
- Too fine: nearly duplicate tickets split apart
You may need to experiment with:
- embedding model
- text preprocessing
- clustering parameters
How to label clusters
After clustering, you can:
- extract common terms from items in a cluster
- ask an LLM to summarize a cluster from representative examples
- manually assign business-friendly labels
Example:
- “Password reset issues”
- “Billing and invoice problems”
- “Search and indexing bugs”
Practical evaluation
To see if clustering is useful:
- Silhouette score for rough separation
- manual review of representative items
- measure whether duplicates or related articles land together
For production, human review is usually the most important.
Example workflow for tickets
- Embed 100,000 historical tickets
- Normalize vectors
- Run HDBSCAN or K-Means
- For each cluster:
- sample 20 tickets
- assign a topic label
- Use cluster labels to:
- route tickets
- detect duplicates
- build knowledge-base article groups
If you want a simple starting point
A good default for many text clustering tasks is:
- embedding model
- normalize vectors
- UMAP for dimensionality reduction
- HDBSCAN for clustering
This combination often works well for semantic grouping of tickets/articles.
If you want, I can give you:
- a Python code example end-to-end, or
- advice on which clustering algorithm to use for your dataset size.
Brands