Prompt
What should I use for GPU autoscaling on model endpoints?
Latest observation
For GPU autoscaling on model endpoints, the best choice depends on where you’re running and whether you need real-time inference or batch/async behavior. Here are the common options, from most practical to more specialized:
Best default choices
1) Kubernetes + KEDA + GPU metrics
Use this if your model endpoints run on Kubernetes.
- KEDA can scale deployments based on queue length, custom metrics, HTTP traffic, or Prometheus metrics.
- Pair it with:
- NVIDIA DCGM exporter for GPU utilization / memory metrics
- Prometheus Adapter or another custom metrics pipeline
- Good when you want:
- scaling based on request rate or queue depth
- custom logic beyond CPU-based scaling
- control over GPU node autoscaling too
Typical setup
- Scale pods with KEDA
- Scale nodes with:
- Cluster Autoscaler or
- Karpenter on AWS, or equivalent on other clouds
2) KServe
Best if you’re serving ML models on Kubernetes and want a higher-level inference platform.
- Supports autoscaling for model serving workloads
- Works well with Knative-style autoscaling
- Better fit for model endpoints than raw Deployments if you want:
- revisioned model rollout
- serverless-style scaling
- inference-specific patterns
3) Ray Serve
Best if your endpoint needs Python-native serving, batching, or more complex inference graphs.
- Has built-in autoscaling
- Useful for:
- dynamic batching
- multi-model pipelines
- custom Python logic
- Good for GPU inference when paired with Kubernetes autoscaling underneath
4) Managed cloud model serving
If you want less ops burden:
- AWS SageMaker endpoints
- GCP Vertex AI endpoints
- Azure ML online endpoints
These can autoscale, but:
- cost and flexibility vary
- GPU scaling may be less transparent
- custom scaling behavior can be limited
What to use in practice
If you’re on Kubernetes:
Use KEDA + Karpenter/Cluster Autoscaler + Prometheus/DCGM metrics
That’s the most flexible and production-friendly setup for GPU autoscaling.
If you want an inference platform:
Use KServe
- especially for standard model endpoints
- if you want simpler serving semantics
If you need advanced Python serving:
Use Ray Serve
- especially for batching, pipelines, or custom logic
Important note about GPU autoscaling
For GPU workloads, scaling on GPU utilization alone is usually not enough:
- GPU utilization can spike or stay high even when latency is fine
- request queue length, inflight requests, or p95 latency are often better signals
A good autoscaling strategy is usually:
- scale pods on queue depth / request concurrency
- scale nodes on pending GPU pod demand
- use GPU metrics as a secondary signal
Rule of thumb
- Simple endpoint on K8s: KEDA
- Model-serving platform: KServe
- Custom Python inference graph: Ray Serve
- Managed service: SageMaker / Vertex / Azure ML
If you tell me your environment:
- cloud provider
- Kubernetes or not
- batch vs online inference
- model framework (PyTorch, vLLM, TGI, Triton, etc.)
…I can recommend a specific stack.