Prompt
Do I need embeddings from an LLM API for semantic search?
Latest observation
Oct 8, 2026
OpenAI APIWeb search: off
Not necessarily.
For semantic search, you need a way to turn text into vectors so you can compare meaning, but those vectors do not have to come from an LLM API.
Common options
-
Hosted embeddings from an API
- Example: OpenAI, Cohere, Voyage, etc.
- Easy to use, usually strong quality.
- Good if you want fast setup and don’t want to manage models.
-
Open-source embedding models
- Example: sentence-transformers, BGE, E5, GTE.
- You can run them locally or on your own server.
- Good for privacy, cost control, or offline use.
-
Traditional search + reranking
- Use keyword search like BM25 first, then rerank with an embedding model or cross-encoder.
- Often a strong practical setup.
Do you need an LLM specifically?
No.
A separate embedding model is typically used. LLMs generate text, while embedding models generate vectors for similarity search.
When API embeddings are useful
- You want high-quality embeddings without model management
- You have moderate usage and can pay per request
- You want to prototype quickly
When local embeddings are useful
- You need lower cost at scale
- You have privacy or compliance requirements
- You want more control over latency and infrastructure
If you want, I can recommend:
- a good API-based semantic search stack, or
- a fully local/open-source stack.