Prompt
How do I use embeddings for recommendation workflows?
Latest observation
Embeddings are a great fit for recommendation systems because they turn users, items, and context into vectors in the same space, so “similar” things are close together. That makes retrieval and ranking much easier.
Core idea
You learn or obtain vector representations for:
- Items: products, movies, articles, songs
- Users: either as a profile vector or a vector built from past interactions
- Context: session, device, time, query, etc.
Then you recommend items whose embeddings are closest to the user/context embedding.
Common recommendation workflows using embeddings
1. Candidate generation / retrieval
Use embeddings to quickly find a small set of promising items from a large catalog.
How it works:
- Compute a user embedding from recent behavior
- Search an ANN index over item embeddings
- Return top-K nearest items
Typical use:
- First stage in a two-stage recommender
- High recall, fast retrieval
Similarity metrics:
- Cosine similarity
- Dot product
- Euclidean distance
2. Personalized ranking
Use embeddings as input features to a ranking model.
How it works:
- Create features from user/item embeddings
- Feed them into a model such as:
- logistic regression
- gradient boosted trees
- neural ranker
- Predict click / purchase / watch probability
Typical use:
- Re-rank the candidates from retrieval
- Optimize precision and business metrics
3. Content-based recommendation
If you have item metadata like text, images, or audio, embeddings can capture semantic similarity.
Examples:
- Recommend similar articles using text embeddings
- Recommend similar products using image embeddings
- Recommend songs with similar audio embeddings
Good for:
- Cold-start items
- Sparse interaction data
4. Collaborative filtering with embeddings
Learn user and item embeddings directly from interaction data.
Example models:
- Matrix factorization
- Two-tower / dual-encoder models
- Neural collaborative filtering
Training signal:
- clicks
- purchases
- watch time
- likes/dislikes
- implicit feedback
Typical architecture: two-tower retrieval
This is one of the most common embedding-based recommendation setups.
Training
- User tower: takes user history/profile and outputs user embedding
- Item tower: takes item features and outputs item embedding
- Train so that positive user-item pairs have high similarity
Serving
- Precompute item embeddings offline
- Compute user embedding online
- Retrieve nearest items from vector index
This works well because item embeddings can be indexed efficiently.
Practical steps to build an embedding-based recommender
Step 1: Define the recommendation objective
Decide what “good” means:
- click-through rate
- conversion
- dwell time
- retention
- revenue
Step 2: Choose embeddings
Options:
- Use pre-trained embeddings for text/images
- Train embeddings on interaction data
- Combine both
Step 3: Build representations
- User embedding: average of recent item embeddings, sequence model, or learned tower
- Item embedding: metadata encoder or learned ID embedding
- Context embedding: session info, query, timestamp, etc.
Step 4: Train on interaction pairs
Positive pairs:
- user clicked item
- user purchased item
- user watched item
Negative pairs:
- random items not interacted with
- in-batch negatives
- sampled negatives from popular items or same category
Step 5: Index item embeddings
Use a vector database or ANN library:
- FAISS
- ScaNN
- Milvus
- Pinecone
- Weaviate
Step 6: Retrieve and rank
- Retrieve top-100 or top-1000 candidates via embedding similarity
- Re-rank with a stronger model using more features
Step 7: Evaluate
Offline metrics:
- Recall@K
- NDCG@K
- MAP
- Hit rate
Online metrics:
- CTR
- conversion
- revenue
- retention
Example: simple workflow
Suppose you run an e-commerce site.
- A user views shoes, sneakers, and running gear.
- You average embeddings of those viewed items to create a user vector.
- You search your item embedding index for nearest products.
- The top results are running shoes, socks, and related accessories.
- You feed those candidates into a ranking model to personalize final ordering.
Best practices
Do
- Update item embeddings when catalog changes
- Use recent behavior for user embeddings
- Combine embedding retrieval with business rules
- Re-rank retrieved items
- Monitor drift and feedback loops
Don’t
- Rely only on embeddings for final ranking
- Ignore cold-start users/items
- Forget negative sampling strategy
- Use stale embeddings for fast-changing catalogs
Challenges
- Cold start: new users/items may have few interactions
- Popularity bias: embeddings can over-recommend popular items
- Feedback loops: recommendations shape future training data
- Scalability: large catalogs need ANN search
- Explainability: vector similarity is harder to interpret
Good use cases
Embeddings work especially well when:
- you have lots of interaction data
- items have rich metadata
- you need semantic similarity
- your catalog is large
- you want scalable retrieval
A simple mental model
Think of embeddings as “coordinates” in a recommendation space:
- Users move based on their preferences
- Items sit where relevant audiences are likely to like them
- The closer they are, the more likely they match
If you want, I can also show:
- a minimal Python example,
- a two-tower architecture diagram, or
- how to do this with FAISS / Pinecone / Milvus.