Prompt
What is the best way to serve a custom LLM without running Kubernetes?
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
If you want to serve a custom LLM without Kubernetes, the best approach is usually:
1) Use a simple containerized serving stack on a single VM or a small set of VMs
For most teams, the sweet spot is:
- Docker for packaging
- A purpose-built model server like:
- vLLM — great for high-throughput GPU inference
- Hugging Face TGI — solid production option
- TensorRT-LLM — best if you’re heavily optimizing for NVIDIA GPUs
- llama.cpp — good for CPU or smaller GPU deployments / quantized models
- A lightweight process manager / reverse proxy:
- systemd, Docker Compose, or Nginx/Traefik
- Optional autoscaling handled outside Kubernetes:
- cloud load balancer + VM autoscaling group, or just manual horizontal scaling
This is the most common “production without K8s” pattern.
2) Choose the serving engine based on your model and hardware
Best general choice: vLLM
Use this if:
- you have an open-weight transformer model
- you want high throughput and lower latency
- you’re serving multiple concurrent requests
- you have NVIDIA GPUs
Why it’s often the default:
- continuous batching
- efficient KV cache management
- OpenAI-compatible server mode
- easy Docker deployment
Best for Hugging Face ecosystem: TGI
Use this if:
- you want a polished, widely used production server
- you prefer Hugging Face-native tooling
- you want simpler deployment with strong defaults
Best for maximum NVIDIA optimization: TensorRT-LLM
Use this if:
- you’re willing to invest in optimization work
- latency and throughput matter a lot
- you’re deploying on NVIDIA GPUs only
Best for edge/CPU/quantized deployment: llama.cpp
Use this if:
- the model is small enough or quantized
- you want CPU inference or lightweight GPU use
- you want very simple operational overhead
3) Run it behind a minimal production stack
A practical setup on a single server:
- Nginx or Traefik as reverse proxy
- vLLM/TGI in Docker
- systemd or Docker restart policies for resilience
- Prometheus + Grafana for monitoring
- OpenTelemetry / structured logs if you need tracing
This gives you:
- TLS termination
- request routing
- health checks
- restarts on failure
- observability
4) If you need scaling, use VMs and a load balancer instead of Kubernetes
For multiple replicas:
- run identical model-serving containers on separate GPU VMs
- put them behind:
- AWS ALB/NLB, GCP Load Balancer, Azure Load Balancer, or
- Nginx/HAProxy
- scale using:
- cloud auto scaling groups
- custom scripts
- Terraform + your CI/CD pipeline
This is much simpler than Kubernetes if your workload is straightforward.
5) Recommended “best overall” option
If you want a single answer:
Use Docker + vLLM on a GPU VM, fronted by Nginx or a cloud load balancer
That setup is usually the best balance of:
- simplicity
- performance
- ease of deployment
- production readiness
6) What not to do
Avoid:
- running the model directly on the host without isolation
- building a custom inference server unless you have a strong reason
- over-engineering with orchestration tools if you only need one or two GPUs
- using Kubernetes just for “best practice” if your scale doesn’t justify it
7) A good decision rule
- Single model, one or a few GPUs, moderate traffic → Docker + vLLM/TGI
- Heavy optimization needs → TensorRT-LLM
- Small model / CPU / quantized edge use → llama.cpp
- Need easy scaling across many nodes → VMs + load balancer, still no Kubernetes required
If you want, I can also give you:
- a reference architecture,
- a Docker Compose example, or
- a comparison of vLLM vs TGI vs llama.cpp for your model.