Prompt

What should I use for embeddings if my team wants to self-host?

Technology / Databases1 observationLast seen Oct 7, 2026

Latest observation

Oct 7, 2026OpenAI APIWeb search: off

If your team wants to self-host embeddings, the best choice usually depends on your constraints, but the safest defaults are:

Good self-hosted options

1. bge-small / bge-base / bge-large

From BAAI. These are very popular for self-hosted semantic search and RAG.

  • Pros: strong retrieval quality, widely used, easy to deploy
  • Cons: slightly more setup if you want best performance
  • Best for: general-purpose document search and RAG

Common picks:

  • bge-small-en-v1.5 — fast, cheap, good baseline
  • bge-base-en-v1.5 — strong balance
  • bge-large-en-v1.5 — better quality, more compute

2. E5 models

Also very solid for retrieval.

  • Pros: high quality, especially for search tasks
  • Cons: requires using the right query/document formatting
  • Best for: search, retrieval, multilingual use cases

Examples:

  • e5-small-v2
  • e5-base-v2
  • multilingual-e5-base

3. Jina embeddings

Good if you want a modern option with strong multilingual support.

  • Pros: good quality, convenient
  • Cons: less “standard” than BGE/E5 in some teams
  • Best for: multilingual or mixed workloads

4. Sentence Transformers models

If you want lots of flexibility and easy self-hosting, this ecosystem is very practical.

  • Pros: huge model catalog, simple deployment
  • Cons: quality varies by model
  • Best for: prototyping and custom workloads

If you want a single recommendation

  • English-only general RAG: bge-base-en-v1.5
  • Need faster/cheaper: bge-small-en-v1.5
  • Multilingual: multilingual-e5-base or a multilingual BGE model
  • Highest quality with more compute: bge-large-en-v1.5

Self-hosting stack suggestions

You can serve embeddings with:

  • SentenceTransformers + FastAPI
  • vLLM if using supported embedding models
  • Hugging Face TEI (Text Embeddings Inference) for a production-ready server
  • ONNX / TensorRT if you want lower latency

What to optimize for

When choosing, ask:

  1. Language coverage — English only or multilingual?
  2. Latency — do you need fast CPU inference?
  3. Accuracy — retrieval quality vs cost
  4. Context size / input length — long docs may matter
  5. GPU availability — big impact on model choice

Practical default

If you’re unsure, start with:

bge-base-en-v1.5 via Hugging Face TEI

It’s a strong, self-hostable baseline that works well for most RAG/search use cases.

If you want, I can also give you:

  • a recommended model by use case
  • a CPU-only self-hosting setup
  • or a Docker example for embedding inference

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.