Prompt
Need GPU autoscaling for inference
Latest observation
Oct 9, 2026
OpenAI APIWeb search: off
Yes — GPU autoscaling for inference is doable, but the “right” setup depends a lot on your serving stack and workload pattern.
Common patterns
1) Kubernetes + HPA/KEDA + GPU nodes
Best if you already run on Kubernetes.
- Pods serve inference
- Horizontal Pod Autoscaler (HPA) or KEDA scales replicas based on:
- request rate
- queue depth
- custom metrics like latency or GPU utilization
- Cluster Autoscaler / Karpenter / Node Autoscaler adds/removes GPU nodes when pods can’t be scheduled
Good for
- TensorRT, vLLM, TGI, Triton, custom PyTorch/TF services
- predictable latency targets
- multi-model or multi-tenant deployments
Caveat
- scaling GPU nodes is slower than scaling pods, so keep some warm capacity.
2) Managed inference platforms
If you want the least ops:
- AWS SageMaker endpoints
- Google Vertex AI
- Azure ML online endpoints
These can autoscale instances and handle much of the infrastructure complexity.
Good for
- simpler operations
- production ML teams without full platform engineering support
Caveat
- less control
- can be more expensive or less flexible for advanced serving patterns
3) Queue-based autoscaling
Very common for inference that is not strictly real-time.
- Requests go into a queue
- Workers consume jobs on GPU
- Scale based on:
- queue length
- age of oldest message
- in-flight requests
Good for
- batch inference
- asynchronous user workflows
- spiky traffic
What to scale on
GPU utilization alone is often not enough. Better signals:
- request queue depth
- p95 latency
- tokens/sec or requests/sec
- concurrent in-flight requests
- GPU memory pressure
- CPU saturation if preprocessing is heavy
For LLMs, scaling on:
- active sequences
- waiting queue length
- token generation throughput
is often more useful than raw GPU utilization.
Practical recommendation
If you’re on Kubernetes:
- Run your inference server in pods
- Use HPA/KEDA on queue depth or custom metrics
- Use Cluster Autoscaler or Karpenter for GPU nodes
- Keep min replicas / min nodes to avoid cold-start pain
- Use model warmup and graceful draining
Important gotchas
- GPU nodes take time to provision
- Model load time can dominate scaling latency
- Memory fragmentation can reduce effective packing
- Batching often improves GPU efficiency more than scaling
- Cold starts can hurt user-facing inference badly
If you want, I can give you one of these next:
- a Kubernetes architecture for GPU autoscaling,
- a KEDA + GPU example YAML,
- a cloud-specific setup for AWS/GCP/Azure,
- or guidance for LLM inference autoscaling specifically.