Prompt
How do I build a vector search pipeline for support tickets?
Latest observation
Here’s a practical blueprint for building a vector search pipeline for support tickets.
1) Define the use case
Common support-ticket vector search use cases:
- Duplicate detection: find similar prior tickets
- Agent assist: retrieve relevant fixes, KB articles, and past resolutions
- Auto-routing: match incoming tickets to teams/categories
- Answer suggestion: surface likely responses based on similar issues
Be clear on:
- What should be searched: ticket subject, description, comments, resolution, KB docs
- What should be returned: similar tickets, suggested answers, or both
- Latency needs: real-time vs batch
- Security/compliance: PII, customer data, access control
2) Prepare the data
Support tickets are usually messy, so cleaning matters.
Typical fields
ticket_idsubjectdescriptioncommentsresolutionstatusprioritycategorycreated_atupdated_atcustomer/orgassignee/team
Preprocessing steps
- Remove or mask PII if needed
- Normalize text:
- strip HTML
- remove signatures
- collapse repeated whitespace
- Optionally enrich:
- append product/version
- include resolution summary
- include tags/category
- Split long tickets into chunks if necessary
Example text representation
A ticket can be embedded as one combined text blob:
Subject: Cannot login after password reset
Description: User receives "invalid token" after resetting password...
Category: Authentication
Resolution: Cleared stale session and forced token refresh
This often works better than embedding only the subject.
3) Choose an embedding model
Use a model that turns ticket text into vectors.
Good options
- Hosted APIs: OpenAI text-embedding models, Cohere, etc.
- Open-source: sentence-transformers models, e5, bge, etc.
Selection criteria
- Quality on short support text
- Cost
- Latency
- Domain adaptation support
- Multilingual support if needed
Best practice
Start with a strong general-purpose embedding model, then evaluate on your ticket data. If needed, fine-tune later.
4) Build the indexing pipeline
Your pipeline typically has these stages:
A. Ingest
Pull tickets from:
- Zendesk
- Salesforce Service Cloud
- Freshdesk
- Jira Service Management
- Custom databases
B. Clean + transform
Create the final text to embed. Example:
- subject
- description
- last resolution/comment
- key metadata
C. Embed
Convert each ticket/document into a vector.
D. Store in a vector database
Options:
- Pinecone
- Weaviate
- Milvus
- Qdrant
- Elasticsearch/OpenSearch vector search
- PostgreSQL + pgvector for simpler setups
Store:
- vector
- ticket metadata
- original text or reference ID
E. Update strategy
- New ticket arrives → embed and upsert
- Ticket updated/resolved → re-embed if content changed
- Periodic reindex if model changes
5) Design the retrieval flow
For a new support ticket:
Step 1: Build query text
Use the incoming ticket’s:
- subject
- description
- initial customer message
- optional metadata
Step 2: Embed query
Generate a vector for the incoming ticket.
Step 3: Retrieve candidates
Run vector similarity search against:
- past tickets
- resolution articles
- internal runbooks
Step 4: Rerank
Use a reranker or cross-encoder to improve precision on the top results.
Step 5: Return results
Show:
- similar tickets
- likely resolution
- confidence score
- matched category/team
6) Use hybrid search for better results
Pure vector search is often not enough for support tickets.
Combine:
- Keyword/BM25 search for exact matches like error codes
- Vector search for semantic similarity
This helps when:
- users paste stack traces
- product names, error codes, or IDs matter
- tickets use exact terminology
A common pattern:
- Use BM25 and vector search in parallel
- Merge results
- Rerank combined candidates
7) Add metadata filters
Support data often needs filtering:
- product line
- language
- region
- customer tier
- date range
- status
- access permissions
Example:
- Search only within the same product
- Exclude resolved tickets older than 2 years
- Restrict customer-specific tickets by tenant
Metadata filtering improves relevance and compliance.
8) Evaluate quality
Measure whether the system helps support agents.
Offline metrics
- Precision@k
- Recall@k
- MRR
- nDCG
Human evaluation
Have agents judge:
- Is the retrieved ticket actually similar?
- Is the suggested resolution useful?
- Is the ranking correct?
Build a gold set
Create labeled pairs:
- query ticket → similar past tickets
- query ticket → correct resolution article
This is essential for tuning chunking, embedding choice, and reranking.
9) Production considerations
Access control
Some tickets may be restricted. Enforce permissions at retrieval time.
Freshness
Index new tickets quickly so agents can search recent cases.
Observability
Track:
- query latency
- top-k hit rate
- click-through on results
- resolution adoption
Feedback loop
Let agents mark results as:
- helpful
- not helpful
- duplicate
- wrong team
Use this feedback to improve ranking.
10) Example architecture
A simple architecture:
-
Ticketing system
New ticket created -
ETL job / stream consumer
Cleans ticket text and metadata -
Embedding service
Generates vector -
Vector DB
Stores vector + metadata -
Search API
For incoming ticket:- embed query
- search vector DB
- apply metadata filters
- optionally hybrid search
- rerank results
-
Agent UI
Displays similar tickets and suggestions
11) Minimal implementation sketch
Pseudocode:
# indexing
for ticket in tickets:
text = build_ticket_text(ticket)
vector = embed(text)
vector_db.upsert(
id=ticket["ticket_id"],
vector=vector,
metadata={
"category": ticket["category"],
"product": ticket["product"],
"status": ticket["status"],
},
text=text
)
# search
query_text = build_ticket_text(new_ticket)
query_vector = embed(query_text)
results = vector_db.search(
vector=query_vector,
top_k=10,
filters={"product": new_ticket["product"]}
)
reranked = rerank(query_text, results)
return reranked[:5]
12) Recommended starting stack
A practical starter stack:
- Embeddings: OpenAI or a sentence-transformers model
- Vector store: pgvector, Qdrant, or Pinecone
- Keyword search: Elasticsearch/OpenSearch or Postgres full-text search
- Reranker: cross-encoder or hosted rerank API
- Orchestration: Python + FastAPI + background jobs
If you want simple and cheap:
- Postgres + pgvector + full-text search
If you want scalable and dedicated:
- Qdrant / Pinecone / Weaviate + reranker
13) Common pitfalls
- Embedding too little text: subject alone is usually not enough
- Not masking PII
- No reranking
- Ignoring exact-match error codes
- Poor chunking of long tickets
- Not filtering by product/team/access
- Not evaluating with real support data
If you want, I can also give you:
- a reference architecture diagram,
- a Python example using pgvector or Qdrant, or
- a production-ready design for Zendesk/Jira support tickets.