Prompt

I'm building product recommendations with embeddings, what should I use?

Technology / Databases1 observationLast seen Oct 7, 2026

Latest observation

Oct 7, 2026OpenAI APIWeb search: off

For product recommendations with embeddings, the best choice usually depends on what you’re recommending from and what data you have. In practice, most teams use a hybrid approach:

Good default

Use two towers / dual-encoder embeddings for retrieval:

  • one embedding for the user/query context
  • one embedding for the product/item
  • then retrieve similar items with nearest-neighbor search

This is a strong starting point because it scales well and works for:

  • “users like this”
  • “similar products”
  • “recommended for you”
  • search-to-product matching

What to use for item embeddings

A product embedding should combine:

  • title
  • description
  • category / taxonomy
  • brand
  • attributes (size, color, material, etc.)
  • maybe image embedding if visuals matter

Common options:

  1. Text embeddings for product metadata

    • easiest and often enough to start
    • works well if titles/descriptions are rich
  2. Multimodal embeddings

    • combine text + image
    • better for fashion, furniture, home goods, etc.
  3. Learned collaborative embeddings

    • from user-item interactions
    • good when you have lots of click/purchase data

Best recommendation setup

A strong production architecture is:

  1. Candidate generation

    • use embeddings + ANN index (FAISS, ScaNN, Pinecone, Weaviate, Milvus, etc.)
    • fetch top 100–1000 candidates fast
  2. Ranking

    • use a separate ranker with more features:
      • price
      • availability
      • popularity
      • user affinity
      • recency
      • business rules
    • often a gradient-boosted model or neural ranker

If you have limited data

Start with:

  • pretrained text embeddings for products
  • average user embeddings from clicked/bought items
  • ANN retrieval
  • simple re-ranking

This gives a fast baseline.

If you have lots of interaction data

Train:

  • matrix factorization / collaborative filtering embeddings
  • or a two-tower retrieval model
  • optionally fine-tune with content features

This usually beats pure content embeddings for personalization.

My practical recommendation

If you’re unsure, start with:

  • product text embeddings
  • user embedding = average of recently clicked/purchased product embeddings
  • ANN search for retrieval
  • then add behavioral/collaborative training later

That’s simple, scalable, and usually good enough to launch.

If you want, I can recommend a concrete stack based on your case:

  • e-commerce catalog size
  • whether you have click/purchase data
  • text only vs image + text
  • real-time vs batch recommendations

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.