Prompt
How do I host embeddings and a chat model together?
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
To host embeddings and a chat model together, the usual pattern is:
- Run both models behind one API service
- Route requests by endpoint:
/embeddings→ embedding model/chator/v1/chat/completions→ chat model
- Optionally add a shared vector store so the chat model can use retrieved context
Common deployment options
1) Single server, two model workers
Best if you want simplicity.
- One API layer (FastAPI, Flask, etc.)
- One embedding model process
- One chat model process
- API routes forward to the right worker
Good for:
- Small/medium traffic
- Different hardware requirements for each model
- Easy maintenance
2) One inference server that supports multiple models
If your serving stack supports it, you can load both models into one system.
Examples:
- vLLM: great for chat/completions, not ideal for embeddings unless you use a separate embedding model service
- TGI (Text Generation Inference): primarily generation
- Ollama: can host multiple models locally, including embedding and chat models
- LM Studio: local dev, not production-focused
3) Separate microservices
Best for production and scaling.
embeddings-servicechat-serviceretrieval-service/ vector DB- API gateway in front
This is usually the cleanest architecture.
Recommended architecture
For a RAG-style app:
Client
↓
API Gateway
├── /embeddings → Embedding service
├── /chat → Chat model service
└── /query → Retrieve docs from vector DB → chat model
Flow
- When indexing documents:
- send text to embedding service
- store vectors in vector DB
- When chatting:
- embed the user query
- search vector DB
- send retrieved context + user prompt to chat model
Example with FastAPI
from fastapi import FastAPI
from pydantic import BaseModel
app = FastAPI()
class EmbedRequest(BaseModel):
text: str
class ChatRequest(BaseModel):
messages: list
@app.post("/embeddings")
def embeddings(req: EmbedRequest):
# call embedding model
vector = [0.1, 0.2, 0.3]
return {"embedding": vector}
@app.post("/chat")
def chat(req: ChatRequest):
# call chat model
response = "Hello! How can I help?"
return {"response": response}
Then internally:
/embeddingsuses an embedding pipeline/chatuses your LLM inference backend
Practical notes
Hardware
- Embedding models are usually much smaller
- Chat models need more GPU memory
- You may want:
- embedding model on CPU or smaller GPU
- chat model on GPU
- or separate GPUs
Latency
- Keep embedding requests fast
- Batch embeddings if indexing lots of data
- Cache frequent embeddings if appropriate
API compatibility
If you want OpenAI-style endpoints:
/v1/embeddings/v1/chat/completions
You can implement both in one service, which makes integration easier.
If you want the easiest self-hosted setup
Use:
- Ollama for both chat + embeddings locally
- or FastAPI + two model backends
- plus Qdrant / Chroma / pgvector for vector search
If you want production scaling
Use:
- Separate embedding and chat services
- Kubernetes or Docker Compose
- Vector DB
- Load balancer / API gateway
If you want, I can give you:
- a Docker Compose example,
- a FastAPI reference implementation, or
- a production architecture diagram.