Prompt

How do I use embeddings for recommendation workflows?

Artificial Intelligence / AI Search1 observationLast seen Oct 10, 2026

Latest observation

Oct 10, 2026OpenAI APIWeb search: off

Embeddings are a great fit for recommendation systems because they turn users, items, and context into vectors in the same space, so “similar” things are close together. That makes retrieval and ranking much easier.

Core idea

You learn or obtain vector representations for:

  • Items: products, movies, articles, songs
  • Users: either as a profile vector or a vector built from past interactions
  • Context: session, device, time, query, etc.

Then you recommend items whose embeddings are closest to the user/context embedding.


Common recommendation workflows using embeddings

1. Candidate generation / retrieval

Use embeddings to quickly find a small set of promising items from a large catalog.

How it works:

  • Compute a user embedding from recent behavior
  • Search an ANN index over item embeddings
  • Return top-K nearest items

Typical use:

  • First stage in a two-stage recommender
  • High recall, fast retrieval

Similarity metrics:

  • Cosine similarity
  • Dot product
  • Euclidean distance

2. Personalized ranking

Use embeddings as input features to a ranking model.

How it works:

  • Create features from user/item embeddings
  • Feed them into a model such as:
    • logistic regression
    • gradient boosted trees
    • neural ranker
  • Predict click / purchase / watch probability

Typical use:

  • Re-rank the candidates from retrieval
  • Optimize precision and business metrics

3. Content-based recommendation

If you have item metadata like text, images, or audio, embeddings can capture semantic similarity.

Examples:

  • Recommend similar articles using text embeddings
  • Recommend similar products using image embeddings
  • Recommend songs with similar audio embeddings

Good for:

  • Cold-start items
  • Sparse interaction data

4. Collaborative filtering with embeddings

Learn user and item embeddings directly from interaction data.

Example models:

  • Matrix factorization
  • Two-tower / dual-encoder models
  • Neural collaborative filtering

Training signal:

  • clicks
  • purchases
  • watch time
  • likes/dislikes
  • implicit feedback

Typical architecture: two-tower retrieval

This is one of the most common embedding-based recommendation setups.

Training

  • User tower: takes user history/profile and outputs user embedding
  • Item tower: takes item features and outputs item embedding
  • Train so that positive user-item pairs have high similarity

Serving

  • Precompute item embeddings offline
  • Compute user embedding online
  • Retrieve nearest items from vector index

This works well because item embeddings can be indexed efficiently.


Practical steps to build an embedding-based recommender

Step 1: Define the recommendation objective

Decide what “good” means:

  • click-through rate
  • conversion
  • dwell time
  • retention
  • revenue

Step 2: Choose embeddings

Options:

  • Use pre-trained embeddings for text/images
  • Train embeddings on interaction data
  • Combine both

Step 3: Build representations

  • User embedding: average of recent item embeddings, sequence model, or learned tower
  • Item embedding: metadata encoder or learned ID embedding
  • Context embedding: session info, query, timestamp, etc.

Step 4: Train on interaction pairs

Positive pairs:

  • user clicked item
  • user purchased item
  • user watched item

Negative pairs:

  • random items not interacted with
  • in-batch negatives
  • sampled negatives from popular items or same category

Step 5: Index item embeddings

Use a vector database or ANN library:

  • FAISS
  • ScaNN
  • Milvus
  • Pinecone
  • Weaviate

Step 6: Retrieve and rank

  • Retrieve top-100 or top-1000 candidates via embedding similarity
  • Re-rank with a stronger model using more features

Step 7: Evaluate

Offline metrics:

  • Recall@K
  • NDCG@K
  • MAP
  • Hit rate

Online metrics:

  • CTR
  • conversion
  • revenue
  • retention

Example: simple workflow

Suppose you run an e-commerce site.

  1. A user views shoes, sneakers, and running gear.
  2. You average embeddings of those viewed items to create a user vector.
  3. You search your item embedding index for nearest products.
  4. The top results are running shoes, socks, and related accessories.
  5. You feed those candidates into a ranking model to personalize final ordering.

Best practices

Do

  • Update item embeddings when catalog changes
  • Use recent behavior for user embeddings
  • Combine embedding retrieval with business rules
  • Re-rank retrieved items
  • Monitor drift and feedback loops

Don’t

  • Rely only on embeddings for final ranking
  • Ignore cold-start users/items
  • Forget negative sampling strategy
  • Use stale embeddings for fast-changing catalogs

Challenges

  • Cold start: new users/items may have few interactions
  • Popularity bias: embeddings can over-recommend popular items
  • Feedback loops: recommendations shape future training data
  • Scalability: large catalogs need ANN search
  • Explainability: vector similarity is harder to interpret

Good use cases

Embeddings work especially well when:

  • you have lots of interaction data
  • items have rich metadata
  • you need semantic similarity
  • your catalog is large
  • you want scalable retrieval

A simple mental model

Think of embeddings as “coordinates” in a recommendation space:

  • Users move based on their preferences
  • Items sit where relevant audiences are likely to like them
  • The closer they are, the more likely they match

If you want, I can also show:

  1. a minimal Python example,
  2. a two-tower architecture diagram, or
  3. how to do this with FAISS / Pinecone / Milvus.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.