Prompt

How do I use embeddings for clustering similar tickets or articles?

Artificial Intelligence / AI Search1 observationLast seen Oct 10, 2026

Latest observation

Oct 10, 2026OpenAI APIWeb search: off

Embeddings are a good way to turn tickets or articles into numeric vectors so you can cluster by semantic similarity instead of just keyword overlap.

High-level approach

  1. Prepare the text

    • Combine useful fields into one string.
    • Examples:
      • Support ticket: subject + description + resolution
      • Article: title + summary + body
    • Clean obvious noise, but don’t over-process.
  2. Generate embeddings

    • Use an embedding model to convert each ticket/article into a vector.
    • Similar texts should have nearby vectors.
  3. Optionally reduce dimensionality

    • Embeddings are often 384–3072 dimensions.
    • For clustering, you may use them directly or reduce with PCA/UMAP.
  4. Cluster the vectors

    • Common choices:
      • K-Means if you know roughly how many clusters you want
      • Hierarchical clustering if you want nested groups
      • DBSCAN/HDBSCAN if you want automatic cluster count and outlier handling
  5. Inspect cluster quality

    • Look at sample items from each cluster.
    • Use cluster labels/centroids to summarize themes.

Typical pipeline

1) Create text input

Example for tickets:

Title: Password reset not working
Body: I cannot reset my password. The reset email never arrives.
Category: Login

2) Embed the text

Using an embedding API or library, get a vector for each item.

3) Cluster

For example, with Python + scikit-learn:

from sklearn.cluster import KMeans

X = embeddings  # shape: (n_items, embedding_dim)
kmeans = KMeans(n_clusters=10, random_state=42)
labels = kmeans.fit_predict(X)

4) Review clusters

For each cluster:

  • print top examples
  • find representative items closest to the centroid

Choosing a clustering method

K-Means

Best if:

  • you have a rough sense of cluster count
  • clusters are fairly spherical in embedding space

Pros:

  • fast, simple Cons:
  • requires choosing k

HDBSCAN / DBSCAN

Best if:

  • you don’t know the number of clusters
  • you expect noise/outliers

Pros:

  • finds natural groupings
  • can mark outliers Cons:
  • more sensitive to parameters
  • can struggle if clusters have very different densities

Hierarchical clustering

Best if:

  • you want a tree of topics or subtopics

Pros:

  • interpretable Cons:
  • can be slower on large datasets

Important tips

Use cosine similarity

For text embeddings, cosine similarity is often better than Euclidean distance.

If you use K-Means, many workflows still work well, especially if embeddings are normalized first.

Normalize embeddings

Often helpful:

from sklearn.preprocessing import normalize
X = normalize(embeddings)

Combine multiple fields carefully

For tickets, good inputs might be:

  • subject
  • description
  • tags
  • resolution notes

But avoid including fields that leak labels or add noise unless they help.

Cluster at the right granularity

  • Too coarse: one cluster contains many unrelated issues
  • Too fine: nearly duplicate tickets split apart

You may need to experiment with:

  • embedding model
  • text preprocessing
  • clustering parameters

How to label clusters

After clustering, you can:

  • extract common terms from items in a cluster
  • ask an LLM to summarize a cluster from representative examples
  • manually assign business-friendly labels

Example:

  • “Password reset issues”
  • “Billing and invoice problems”
  • “Search and indexing bugs”

Practical evaluation

To see if clustering is useful:

  • Silhouette score for rough separation
  • manual review of representative items
  • measure whether duplicates or related articles land together

For production, human review is usually the most important.


Example workflow for tickets

  1. Embed 100,000 historical tickets
  2. Normalize vectors
  3. Run HDBSCAN or K-Means
  4. For each cluster:
    • sample 20 tickets
    • assign a topic label
  5. Use cluster labels to:
    • route tickets
    • detect duplicates
    • build knowledge-base article groups

If you want a simple starting point

A good default for many text clustering tasks is:

  • embedding model
  • normalize vectors
  • UMAP for dimensionality reduction
  • HDBSCAN for clustering

This combination often works well for semantic grouping of tickets/articles.

If you want, I can give you:

  1. a Python code example end-to-end, or
  2. advice on which clustering algorithm to use for your dataset size.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.