Prompt
What should I use to host embeddings and generation endpoints?
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
It depends on whether you want speed/cost control, simplicity, or managed convenience. A good default split is:
- Embeddings endpoint: use a smaller, cheaper, high-throughput model on a CPU-friendly or lightly GPU-backed inference service
- Generation endpoint: use a GPU-backed inference service with autoscaling, since generation is much more compute-heavy
Good hosting options
1) Managed cloud inference platforms
Best if you want to move fast.
Examples
- AWS SageMaker / Bedrock
- Google Vertex AI
- Azure AI Foundry / Azure ML
- Hugging Face Inference Endpoints
- Replicate
- Together.ai / Fireworks.ai / Groq / Anyscale (depending on model support)
Pros
- Easy deployment
- Autoscaling
- Monitoring and auth are usually built in
- Less ops work
Cons
- Can get expensive at scale
- Less control over performance tuning
Best for
- Early production
- Small teams
- Teams that don’t want to manage infrastructure
2) Self-hosted on Kubernetes or VMs
Best if you care about cost optimization and control.
Common serving stacks
- vLLM: great for generation throughput
- TGI (Text Generation Inference): solid for LLM serving
- TensorRT-LLM: high performance, more complex
- BentoML / Ray Serve / KServe / Seldon: orchestration layers
- ONNX Runtime / OpenVINO / llama.cpp / TEI for embeddings, depending on model type
Pros
- Full control
- Often cheaper at scale
- Easier to optimize latency/throughput exactly how you want
Cons
- More engineering/ops
- You manage scaling, deployments, observability
Best for
- Mature teams
- High volume
- Strict latency/cost goals
Recommended approach by endpoint type
Embeddings
For embeddings, I’d usually use:
- Hosted API if volume is low/medium and you want simplicity
- Self-hosted embedding model if volume is high or data privacy matters
Good serving choices:
- Sentence-Transformers model served via FastAPI + ONNX Runtime
- TEI (Text Embeddings Inference) if using Hugging Face models
- vLLM is generally more for generation than embeddings
Generation
For generation, I’d usually use:
- vLLM if self-hosting
- TGI if you want a more packaged HF-native option
- A managed LLM provider if you want minimal ops
Practical recommendation
If you’re asking “what should I use?” and want a balanced answer:
- Embeddings: Hugging Face TEI or a small custom service with ONNX Runtime
- Generation: vLLM on GPU instances
- Hosting: Kubernetes if you have DevOps capacity, otherwise a managed endpoint like SageMaker, Vertex AI, or Hugging Face Inference Endpoints
Simple rule of thumb
- Need fastest setup: managed endpoints
- Need best cost at scale: self-host with vLLM/TGI + TEI
- Need least ops: a single managed provider for both
- Need privacy/control: self-host both
If you want, I can recommend a specific stack based on your:
- expected QPS,
- model size,
- latency target,
- cloud provider,
- budget.