Prompt

How do I set up autoscaling for GPU model serving?

Artificial Intelligence / AI Infrastructure1 observationLast seen Oct 10, 2026

Latest observation

Oct 10, 2026OpenAI APIWeb search: off

To set up autoscaling for GPU model serving, you usually need to scale on workload demand rather than just CPU/memory, because GPU inference is often constrained by:

  • request rate / queue depth
  • GPU utilization
  • batching latency
  • active concurrent requests
  • model load time / warmup time

Here’s a practical setup pattern that works well.


1) Pick the right serving architecture

Common options:

  • Kubernetes + GPU nodes
    Most flexible and common for production.
  • Managed inference platforms
    Easier operationally, less control.
  • Triton / vLLM / TGI / TorchServe behind a service layer
    Good if you need optimized batching and streaming.

For most GPU autoscaling setups on Kubernetes, you want:

  • a Deployment or StatefulSet for inference pods
  • GPU node pool in your cluster
  • Horizontal Pod Autoscaler (HPA) or KEDA
  • Cluster Autoscaler or equivalent to add/remove GPU nodes

2) Use metrics that reflect inference load

CPU-based scaling is usually not enough. Better signals:

Good autoscaling metrics

  • Requests per second
  • In-flight requests
  • Queue length
  • P95/P99 latency
  • GPU utilization
  • GPU memory usage
  • Tokens/sec for LLMs
  • Time spent waiting in request queue

Best practice

Scale on a metric tied to serving saturation, like:

  • queue depth per replica
  • concurrent requests per replica
  • GPU utilization averaged over a short window

If you use LLM serving, queue length + latency is often more reliable than raw GPU utilization because batching can make GPU usage look high even before saturation.


3) Make the service autoscaling-friendly

Before autoscaling, make sure your inference server can scale horizontally:

  • Each pod loads the model independently, or uses a shared weight cache if supported
  • Use readiness probes so traffic only reaches warm pods
  • Support graceful shutdown so in-flight requests finish
  • Prefer dynamic batching if supported
  • Ensure model startup time is known, because GPU pods can take a while to initialize

If startup is slow, consider:

  • min replicas > 0
  • pre-warmed pods
  • scale-up buffers
  • image/model caching on nodes

4) Set up Kubernetes autoscaling components

A. HPA for pods

Use HPA to scale the number of serving pods based on custom metrics.

Typical config uses:

  • custom metrics from Prometheus
  • metrics adapter like Prometheus Adapter
  • or KEDA for queue-based/event-driven scaling

B. Cluster Autoscaler for GPU nodes

If HPA adds pods but there are no GPU nodes available, the cluster must scale the node pool too.

Requirements:

  • GPU node pool configured separately
  • Cluster Autoscaler enabled
  • node pool labels/taints aligned with pod node selectors/tolerations

5) Example scaling pattern

A common production flow:

  1. Traffic increases
  2. Queue depth / latency rises
  3. HPA increases replicas from 2 → 6
  4. Pods request GPUs
  5. Cluster Autoscaler adds GPU nodes
  6. New pods start once nodes are ready

To avoid request spikes causing user-visible delays:

  • keep a minimum replica count
  • set scale-up stabilization
  • use pod anti-affinity to spread replicas across nodes
  • add buffer capacity if traffic is bursty

6) Example Kubernetes resources

GPU deployment snippet

apiVersion: apps/v1
kind: Deployment
metadata:
  name: model-server
spec:
  replicas: 2
  selector:
    matchLabels:
      app: model-server
  template:
    metadata:
      labels:
        app: model-server
    spec:
      containers:
      - name: server
        image: your-model-server:latest
        resources:
          limits:
            nvidia.com/gpu: 1

HPA using custom metrics

Example conceptually scales on a custom metric like requests_per_pod:

apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: model-server-hpa
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: model-server
  minReplicas: 2
  maxReplicas: 10
  metrics:
  - type: Pods
    pods:
      metric:
        name: inference_requests_per_second
      target:
        type: AverageValue
        averageValue: "20"

In practice, you’d expose that metric via:

  • Prometheus
  • OpenTelemetry
  • KEDA scaler
  • cloud provider metrics

7) Consider KEDA for queue-based scaling

KEDA is often a very good fit when requests land in a queue first.

Examples:

  • Kafka
  • RabbitMQ
  • Azure Queue
  • AWS SQS
  • Redis streams

Why it helps:

  • scales based on actual backlog
  • works well for async inference
  • can reduce over-scaling during short spikes

If your system is synchronous HTTP inference, HPA with custom metrics is usually better.


8) GPU-specific scaling tips

Avoid cold-start pain

GPU pods can be slow to start due to:

  • image pull time
  • model load time
  • CUDA initialization
  • compilation/warmup

Mitigate with:

  • node image cache
  • pre-pulled images
  • smaller images
  • model warmup requests
  • higher min replicas

Scale on replicas, not GPU percentage alone

A single GPU at 60% may still be overloaded if:

  • queue is growing
  • latency is rising
  • memory is near limit

Use batching carefully

Dynamic batching improves throughput but can increase latency if overdone. Tune:

  • max batch size
  • queue delay
  • concurrency per replica

9) Monitoring you should have

At minimum, monitor:

  • request latency p50/p95/p99
  • throughput
  • error rate
  • queue depth
  • GPU utilization
  • GPU memory usage
  • pod startup time
  • model load time
  • number of ready replicas

Use alerts for:

  • sustained high queue depth
  • replicas at max
  • node pool at max
  • GPU OOMs
  • readiness failures

10) Recommended starting policy

A reasonable starting point:

  • minReplicas: 1–2
  • maxReplicas: enough to handle peak traffic
  • scale up quickly, scale down slowly
  • target GPU utilization: 60–75%
  • target queue length per pod: low and bounded
  • warmup/readiness delay: ensure pod is ready only after model is loaded

For LLM serving, start with:

  • scale on in-flight requests
  • use batching
  • keep at least 1 warm replica
  • scale node pool with a buffer

11) If you want the simplest reliable setup

Use:

  • Kubernetes
  • GPU node pool
  • vLLM/Triton/TGI
  • Prometheus metrics
  • HPA or KEDA
  • Cluster Autoscaler

This is the most common “works in production” stack.


If you want, I can give you:

  1. a Kubernetes YAML example for GPU HPA,
  2. a KEDA queue-scaling example, or
  3. a cloud-specific setup for AWS EKS, GKE, or AKS.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.