Prompt
How do I orchestrate GPUs for real-time inference?
Latest observation
To orchestrate GPUs for real-time inference, think in terms of latency, placement, and isolation rather than just raw throughput. The main goal is to keep models warm, requests routed to the right GPU, and tail latency under control.
1) Start with the right serving architecture
A good GPU inference setup usually has:
- Model server per GPU or per GPU slice
- Examples: NVIDIA Triton, vLLM, TensorRT-LLM, TorchServe, BentoML, Ray Serve
- A lightweight router / gateway
- Routes requests to the best model instance based on latency, load, tenant, or model version
- Autoscaling layer
- Scales replicas or GPU nodes based on queue depth, GPU utilization, or request rate
2) Use GPU-aware scheduling
Your orchestrator needs to know:
- Which nodes have free GPUs
- Whether a GPU already has the model loaded
- Memory capacity per model
- Whether the workload is latency-sensitive or batchable
Common approaches:
- Kubernetes + GPU device plugin
- Kubernetes + Node labels / taints / tolerations
- KubeRay / Ray Serve for placement-aware inference
- NVIDIA DCGM + Prometheus for GPU metrics
If you use Kubernetes:
- Install the NVIDIA GPU Operator
- Use node selectors or affinity to place inference pods on GPU nodes
- Set resource requests/limits for
nvidia.com/gpu - Consider topology-aware scheduling if multi-GPU models need NVLink-local placement
3) Keep models warm
For real-time inference, cold starts hurt.
Best practices:
- Preload model weights at container startup
- Keep a minimum number of always-on replicas
- Avoid scale-to-zero for latency-critical models unless your SLA allows it
- Use probes that check model readiness, not just container health
4) Optimize for tail latency
Important for real-time:
- Use micro-batching carefully; it improves throughput but can increase latency
- Set tight max batch delay windows
- Prefer smaller, faster models or distillation/quantization
- Use FP16 / BF16 / INT8 where acceptable
- Pin CPU threads and enable NUMA-aware placement if the host matters
5) Separate traffic classes
Don’t mix everything on the same serving path.
Recommended:
- Interactive/real-time
- Dedicated replicas, stricter latency SLOs
- Offline/batch
- Can use opportunistic spare GPU capacity
- Experimental / canary
- Small traffic fraction, separate deployment
This avoids batch jobs causing latency spikes for live traffic.
6) Choose a scaling strategy
Scale on metrics that correlate with user experience:
- Request queue length
- P95/P99 latency
- GPU memory pressure
- Token generation rate for LLMs
- Active sessions / concurrent streams
For LLMs, GPU utilization alone is not enough. A GPU can be “busy” while still under-serving because of fragmentation or long-running sequences.
7) Use the right model-serving primitives
Depending on the model type:
For LLMs
Use:
- vLLM
- TensorRT-LLM
- TGI (Text Generation Inference)
Why:
- Continuous batching
- KV-cache management
- Better token-level throughput/latency tradeoff
For CV / classic DL
Use:
- Triton Inference Server
- ONNX Runtime with TensorRT / CUDA EP
- TensorRT directly for maximum performance
8) Multi-tenancy and isolation
If multiple teams or models share GPUs:
- Use MIG on supported NVIDIA GPUs
- Use time-slicing only if latency jitter is acceptable
- Enforce per-tenant quotas and priorities
- Use separate namespaces and resource limits in Kubernetes
For strict real-time SLA, dedicated GPUs or MIG slices are usually better than oversubscribed sharing.
9) Add observability
Track:
- GPU utilization
- GPU memory used
- Per-model latency p50/p95/p99
- Queue depth
- Request rates
- Error rates
- Cold starts
- KV-cache usage for LLMs
Useful tools:
- Prometheus + Grafana
- NVIDIA DCGM exporter
- OpenTelemetry for request tracing
10) A practical Kubernetes setup
Typical deployment flow:
- Provision GPU nodes
- Install NVIDIA drivers + GPU Operator
- Deploy inference server with GPU requests
- Add a service mesh or API gateway
- Scale replicas using HPA/KEDA
- Monitor latency and GPU metrics
- Introduce canary rollout for model updates
Example conceptually:
deployment: llm-serverresources: limits: nvidia.com/gpu: 1service: llm-servicehpa: based on queue depth and latencygateway: routes requests by model name/version
11) If you need low-latency across multiple GPUs
Consider:
- Model parallelism only when necessary; it adds communication overhead
- Tensor parallelism for large models
- Place communicating GPUs on the same PCIe root/NVLink domain
- Use NCCL-aware placement and avoid cross-node sharding unless unavoidable
12) Recommended stack by use case
Best general-purpose
- Kubernetes
- NVIDIA GPU Operator
- Triton Inference Server
- Prometheus/Grafana
- KEDA or HPA
Best for LLM real-time
- Kubernetes or bare metal
- vLLM or TensorRT-LLM
- API gateway
- Autoscaling on queue depth / token latency
- MIG or dedicated GPUs for isolation
Best for simplest operations
- A small number of dedicated inference nodes
- One model server per node
- Simple load balancer
- Manual scaling plus metrics-based alerts
If you want, I can give you one of these next:
- a Kubernetes reference architecture,
- a vLLM/Triton deployment example, or
- a GPU autoscaling strategy for real-time LLM serving.