Prompt
How do I set up autoscaling for GPU model serving?
Latest observation
To set up autoscaling for GPU model serving, you usually need to scale on workload demand rather than just CPU/memory, because GPU inference is often constrained by:
- request rate / queue depth
- GPU utilization
- batching latency
- active concurrent requests
- model load time / warmup time
Here’s a practical setup pattern that works well.
1) Pick the right serving architecture
Common options:
- Kubernetes + GPU nodes
Most flexible and common for production. - Managed inference platforms
Easier operationally, less control. - Triton / vLLM / TGI / TorchServe behind a service layer
Good if you need optimized batching and streaming.
For most GPU autoscaling setups on Kubernetes, you want:
- a Deployment or StatefulSet for inference pods
- GPU node pool in your cluster
- Horizontal Pod Autoscaler (HPA) or KEDA
- Cluster Autoscaler or equivalent to add/remove GPU nodes
2) Use metrics that reflect inference load
CPU-based scaling is usually not enough. Better signals:
Good autoscaling metrics
- Requests per second
- In-flight requests
- Queue length
- P95/P99 latency
- GPU utilization
- GPU memory usage
- Tokens/sec for LLMs
- Time spent waiting in request queue
Best practice
Scale on a metric tied to serving saturation, like:
- queue depth per replica
- concurrent requests per replica
- GPU utilization averaged over a short window
If you use LLM serving, queue length + latency is often more reliable than raw GPU utilization because batching can make GPU usage look high even before saturation.
3) Make the service autoscaling-friendly
Before autoscaling, make sure your inference server can scale horizontally:
- Each pod loads the model independently, or uses a shared weight cache if supported
- Use readiness probes so traffic only reaches warm pods
- Support graceful shutdown so in-flight requests finish
- Prefer dynamic batching if supported
- Ensure model startup time is known, because GPU pods can take a while to initialize
If startup is slow, consider:
- min replicas > 0
- pre-warmed pods
- scale-up buffers
- image/model caching on nodes
4) Set up Kubernetes autoscaling components
A. HPA for pods
Use HPA to scale the number of serving pods based on custom metrics.
Typical config uses:
- custom metrics from Prometheus
- metrics adapter like Prometheus Adapter
- or KEDA for queue-based/event-driven scaling
B. Cluster Autoscaler for GPU nodes
If HPA adds pods but there are no GPU nodes available, the cluster must scale the node pool too.
Requirements:
- GPU node pool configured separately
- Cluster Autoscaler enabled
- node pool labels/taints aligned with pod node selectors/tolerations
5) Example scaling pattern
A common production flow:
- Traffic increases
- Queue depth / latency rises
- HPA increases replicas from 2 → 6
- Pods request GPUs
- Cluster Autoscaler adds GPU nodes
- New pods start once nodes are ready
To avoid request spikes causing user-visible delays:
- keep a minimum replica count
- set scale-up stabilization
- use pod anti-affinity to spread replicas across nodes
- add buffer capacity if traffic is bursty
6) Example Kubernetes resources
GPU deployment snippet
apiVersion: apps/v1
kind: Deployment
metadata:
name: model-server
spec:
replicas: 2
selector:
matchLabels:
app: model-server
template:
metadata:
labels:
app: model-server
spec:
containers:
- name: server
image: your-model-server:latest
resources:
limits:
nvidia.com/gpu: 1
HPA using custom metrics
Example conceptually scales on a custom metric like requests_per_pod:
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: model-server-hpa
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: model-server
minReplicas: 2
maxReplicas: 10
metrics:
- type: Pods
pods:
metric:
name: inference_requests_per_second
target:
type: AverageValue
averageValue: "20"
In practice, you’d expose that metric via:
- Prometheus
- OpenTelemetry
- KEDA scaler
- cloud provider metrics
7) Consider KEDA for queue-based scaling
KEDA is often a very good fit when requests land in a queue first.
Examples:
- Kafka
- RabbitMQ
- Azure Queue
- AWS SQS
- Redis streams
Why it helps:
- scales based on actual backlog
- works well for async inference
- can reduce over-scaling during short spikes
If your system is synchronous HTTP inference, HPA with custom metrics is usually better.
8) GPU-specific scaling tips
Avoid cold-start pain
GPU pods can be slow to start due to:
- image pull time
- model load time
- CUDA initialization
- compilation/warmup
Mitigate with:
- node image cache
- pre-pulled images
- smaller images
- model warmup requests
- higher min replicas
Scale on replicas, not GPU percentage alone
A single GPU at 60% may still be overloaded if:
- queue is growing
- latency is rising
- memory is near limit
Use batching carefully
Dynamic batching improves throughput but can increase latency if overdone. Tune:
- max batch size
- queue delay
- concurrency per replica
9) Monitoring you should have
At minimum, monitor:
- request latency p50/p95/p99
- throughput
- error rate
- queue depth
- GPU utilization
- GPU memory usage
- pod startup time
- model load time
- number of ready replicas
Use alerts for:
- sustained high queue depth
- replicas at max
- node pool at max
- GPU OOMs
- readiness failures
10) Recommended starting policy
A reasonable starting point:
- minReplicas: 1–2
- maxReplicas: enough to handle peak traffic
- scale up quickly, scale down slowly
- target GPU utilization: 60–75%
- target queue length per pod: low and bounded
- warmup/readiness delay: ensure pod is ready only after model is loaded
For LLM serving, start with:
- scale on in-flight requests
- use batching
- keep at least 1 warm replica
- scale node pool with a buffer
11) If you want the simplest reliable setup
Use:
- Kubernetes
- GPU node pool
- vLLM/Triton/TGI
- Prometheus metrics
- HPA or KEDA
- Cluster Autoscaler
This is the most common “works in production” stack.
If you want, I can give you:
- a Kubernetes YAML example for GPU HPA,
- a KEDA queue-scaling example, or
- a cloud-specific setup for AWS EKS, GKE, or AKS.