Prompt
How do I search similar products using embeddings?
Latest observation
To search for similar products with embeddings, you usually do this:
1) Convert each product into text
Create a text representation that captures the product well, for example:
- title
- brand
- category
- key attributes
- short description
Example:
Nike Air Zoom Pegasus 40 running shoes, men's, black, size 10, breathable mesh, cushioned sole
2) Generate an embedding for each product
Use an embedding model to turn each product text into a vector, e.g. 1536 dimensions.
You do this for:
- every product in your catalog
- the user’s search query
3) Store product embeddings in a vector database
Common options:
- pgvector in PostgreSQL
- Pinecone
- Weaviate
- Milvus
- FAISS
- Elasticsearch / OpenSearch vector search
Store:
- product_id
- embedding vector
- metadata like price, category, availability
4) Embed the search query
If a user searches:
lightweight black running shoes for men
Generate an embedding for that query using the same model.
5) Find nearest neighbors
Compare the query embedding to all product embeddings using a similarity metric:
- cosine similarity
- dot product
- Euclidean distance
Return the top-k most similar products.
Basic flow
Offline indexing
- Build text for each product
- Create embeddings
- Save embeddings in vector DB
Online search
- Convert user query to embedding
- Search nearest vectors
- Optionally filter by metadata
- Rank and return results
Example in Python-like pseudocode
# indexing
for product in products:
text = f"{product['title']} {product['brand']} {product['category']} {product['description']}"
product['embedding'] = embed(text)
vector_db.upsert(product_id=product['id'], vector=product['embedding'], metadata=product)
# search
query = "lightweight black running shoes for men"
query_vec = embed(query)
results = vector_db.search(vector=query_vec, top_k=10)
Good practices
Use rich product text
Include the important attributes that affect similarity:
- category
- brand
- color
- material
- use case
- size/specs
- tags
Normalize product text
Make formatting consistent so embeddings are more reliable.
Combine semantic search with filters
Example:
- semantic similarity for “similar products”
- filters for category = shoes, price < 100, in_stock = true
Re-rank results if needed
You can use a second-stage ranking model or business rules to improve relevance.
Example query strategies
“Search by example”
User clicks a product and wants similar items:
- embed the clicked product
- search nearest neighbors
“Search by text”
User types natural language:
- embed the query
- retrieve matching products
“Hybrid search”
Combine:
- keyword search
- embedding similarity
This often works better than embeddings alone, especially for exact attributes like model numbers or sizes.
A simple product similarity formula
If vectors are normalized, cosine similarity works well:
similarity = dot(query_vector, product_vector)
Higher score = more similar.
If you want, I can also show:
- a full code example with Python + pgvector
- how to do this with OpenAI embeddings
- how to build a “similar products” endpoint in an API