Prompt
What should I use to host an LLM in production?
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
It depends mainly on scale, latency, budget, and how much ops you want to own. For production LLM hosting, the usual options are:
1) Managed LLM APIs
Best if you want the fastest path to production.
- Examples: OpenAI, Anthropic, Google Gemini, Azure OpenAI
- Pros: no infra to manage, strong reliability, easy scaling
- Cons: per-token cost, less control, vendor lock-in, data/compliance constraints
2) Model hosting platforms / inference APIs
Best if you want to use open-source models without managing GPUs.
- Examples: Together AI, Fireworks, Replicate, Hugging Face Inference Endpoints, Baseten, Anyscale
- Pros: simpler than self-hosting, can swap models more easily, often good latency
- Cons: still pay platform margin, less control than self-hosting
3) Self-hosted on cloud GPUs
Best if you need control, custom models, or cost optimization at scale.
- Run on: AWS/GCP/Azure GPU instances, or cheaper GPU clouds like Lambda, CoreWeave, RunPod, Vast
- Serving stacks: vLLM, TensorRT-LLM, TGI, SGLang, Ollama for lighter use
- Pros: maximum control, can be cheapest at high volume, easier compliance in some cases
- Cons: highest ops burden, scaling/monitoring/reliability are on you
4) On-prem / private cloud
Best for strict compliance, data residency, or existing GPU infrastructure.
- Pros: full control, strongest isolation
- Cons: highest complexity, capex, maintenance
My default recommendation
- Startup / MVP: use a managed API first.
- Open-source model + low ops: use a hosted inference platform.
- High volume or strict control: self-host with vLLM or TensorRT-LLM on cloud GPUs.
What to choose by use case
- Chat app / agent prototype: OpenAI or Anthropic
- Need open-source model: Fireworks, Together, Hugging Face Endpoints, Baseten
- Heavy throughput / cost-sensitive: self-host with vLLM on AWS/GCP/CoreWeave
- Enterprise compliance: Azure OpenAI or self-host in your VPC
Good production stack if you self-host
- Serving: vLLM
- Orchestration: Kubernetes or autoscaled VM pool
- Monitoring: Prometheus/Grafana + tracing
- Load balancing: NGINX/Envoy
- Queueing / rate limits: Redis, Kafka, or API gateway
- Model versions: keep a registry and canary deploys
Rule of thumb
- If you’re asking “what should I use?”, the safest answer is: Use a managed LLM API unless you have a clear reason not to.
- If you need open-source hosting with minimal ops: Use a hosted inference platform.
- If you’re doing serious scale or need full control: Use vLLM on cloud GPUs.
If you tell me:
- your expected requests/sec or token volume,
- whether you need open-source or a closed model,
- your latency target, and
- whether compliance matters,
I can recommend a specific hosting setup.