Prompt

What should I use for a cheap embeddings pipeline with acceptable recall?

Artificial Intelligence / AI Search1 observationLast seen Oct 10, 2026

Latest observation

Oct 10, 2026OpenAI APIWeb search: off

If you want cheap embeddings with “good enough” recall, the best choice is usually:

Recommended default

sentence-transformers/all-MiniLM-L6-v2

  • Very cheap / fast
  • Small model (~22M params)
  • Works well for general semantic search
  • Common baseline for production prototypes and budget systems

Why it’s a good fit

  • High throughput on CPU
  • Low memory use
  • Strong quality for its size
  • Easy to deploy locally or on a modest API budget

If you want a bit better recall

BAAI/bge-small-en-v1.5

  • Often better retrieval quality than MiniLM
  • Still relatively cheap
  • Good choice for English-only search

If you can afford slightly more compute, this is often the better “cheap but decent” option.


If multilingual is needed

intfloat/multilingual-e5-small

  • Good for multilingual retrieval
  • Still lightweight
  • Use the right prompting format for E5 models, e.g.:
    • query: "query: ..."
    • document: "passage: ..."

Practical pipeline for cheap + acceptable recall

If cost matters, don’t overspend on embeddings alone. Use:

  1. Small embedding model
    • MiniLM or BGE-small
  2. ANN vector index
    • FAISS, hnswlib, or pgvector
  3. Chunking
    • Use sensible chunk sizes, e.g. 200–500 tokens
  4. Reranking for top results
    • Cheap embeddings + a small cross-encoder reranker on top 20–50 results can improve quality a lot

This usually gives much better recall/precision tradeoff than using a large embedding model everywhere.


If you want the absolute cheapest acceptable setup

  • Model: all-MiniLM-L6-v2
  • Index: FAISS HNSW or IVF
  • Reranker: optional, only if quality matters
  • Chunk size: 300–400 tokens with overlap

Quick rule of thumb

  • Lowest cost: all-MiniLM-L6-v2
  • Best cheap quality: bge-small-en-v1.5
  • Multilingual cheap: multilingual-e5-small

If you tell me:

  • your language(s),
  • data size,
  • latency budget,
  • and whether this is for semantic search / RAG / recommendations,

I can suggest a more specific stack.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.