Prompt
How do I host embeddings and chat models in the same serving layer?
Latest observation
You can host embeddings and chat/completion models in the same serving layer, but the cleanest way is to treat them as different model types behind one router/API, not as one identical inference path.
Practical patterns
1) One API gateway, multiple model backends
Expose a single service endpoint, then route requests by:
- task:
embedvschat - model name
- tenant / SLA / GPU pool
Example:
POST /v1/embeddings→ embedding modelPOST /v1/chat/completions→ chat model
This is the simplest and most common approach.
2) One inference platform with shared serving runtime
Use a serving stack that supports multiple model endpoints on the same cluster:
- vLLM
- TGI
- Triton
- Ray Serve
- BentoML
- KServe/Seldon on Kubernetes
Then deploy:
- one or more embedding model replicas
- one or more chat model replicas
- a thin router or gateway in front
3) Model multiplexing on the same GPU nodes
If your traffic is moderate, you can colocate embeddings and chat models on the same nodes:
- embeddings are usually shorter, cheaper, and higher-throughput
- chat is longer-running and more memory intensive
To avoid contention:
- keep separate queues
- set concurrency limits per model
- use different autoscaling policies
- pin large chat models to bigger GPUs, embeddings to smaller GPUs if possible
Recommended architecture
Frontend
- API gateway / load balancer
- auth, rate limiting, request normalization
Router
Decides:
- which model to use
- which backend pool to hit
- whether to batch requests
Backends
- Embedding service: optimized for batch throughput and low latency
- Chat service: optimized for generation latency and token streaming
Shared infrastructure
- metrics/logging/tracing
- model registry
- vector DB for embeddings
- prompt/template store for chat
Important design differences
Embeddings
- usually fixed-length vector output
- benefits from batching
- latency-sensitive but less compute-heavy
- often best served with encoder-only models
Chat models
- autoregressive generation
- token-by-token streaming
- benefits from:
- paged attention / KV-cache management
- batching across concurrent requests
- separate priority queues
Because of that, you usually should not force both through exactly the same execution path internally.
A good implementation strategy
If you want simplicity
Use two model endpoints behind one API server:
/embed/chat
If you want efficiency
Use a shared serving cluster with:
- a router service
- separate worker pools
- shared GPU nodes only when utilization is low enough
If you want production robustness
Use:
- Kubernetes
- separate deployments for embedding/chat
- HPA based on queue length, GPU utilization, and p95 latency
- canary releases per model
Example request routing logic
if request.type == "embedding":
route to embedding pool
elif request.type == "chat":
route to chat pool
else:
reject
You can also route by model:
text-embedding-3-large→ embedding backendllama-3-70b-instruct→ chat backend
Tips to avoid performance problems
- Don’t batch chat and embedding requests together.
- Keep separate autoscaling.
- Use separate queues to prevent long chat requests from blocking embeddings.
- Prefer smaller embedding models on cheaper hardware.
- Stream chat responses; embeddings should return immediately.
- Cache embeddings for repeated inputs.
- Limit max context length for chat to reduce tail latency.
If you mean “one server binary”
You can absolutely build one server that:
- loads an embedding model
- loads a chat model
- exposes both APIs
But be careful:
- memory footprint can get large
- GPU fragmentation can hurt throughput
- different decoding paths make tuning harder
This is best for:
- prototypes
- low-traffic internal services
- edge deployments with one or two models
Rule of thumb
- One API, multiple model services = best balance
- One cluster, separate deployments = best production setup
- One process for everything = easiest, but least scalable
If you want, I can give you:
- a reference architecture diagram,
- a Kubernetes deployment pattern, or
- an example using vLLM / Triton / Ray Serve.