Prompt
Why do I need separate vendors for embeddings and chat?
Latest observation
Oct 9, 2026
OpenAI APIWeb search: off
You don’t always need separate vendors for embeddings and chat — but many teams end up using different services because the two jobs have different strengths and tradeoffs.
Why they’re often separate
-
Different model types
- Embeddings turn text into vectors for search, retrieval, clustering, and similarity.
- Chat models generate natural language responses.
- These are different workloads, so vendors often optimize different models for each.
-
Cost and performance
- Embedding generation is usually high-volume and needs to be cheap, fast, and consistent.
- Chat is lower-volume but needs reasoning, instruction-following, and better generation quality.
- One vendor may be best for one but not the other.
-
Quality differences
- A vendor might have an excellent chat model but mediocre embeddings, or vice versa.
- For retrieval-augmented generation (RAG), embedding quality can matter a lot for search relevance.
-
Operational flexibility
- Some teams want to swap chat models without redoing their retrieval pipeline.
- Others want one vendor for simpler billing, auth, logging, and compliance.
Why you might use the same vendor
- Simpler integration
- Unified billing and monitoring
- Easier governance and security
- Potentially better compatibility between embedding and generation stack
Bottom line
You need separate vendors only if it helps your application. Many applications can use one vendor for both. Teams split vendors when they want the best quality, lowest cost, or more flexibility for each part of the system.
If you want, I can also explain this in terms of a RAG pipeline and where embeddings vs chat fit.