Prompt
I'm building an internal AI platform—how do I manage GPU scheduling and autoscaling?
Latest observation
Managing GPU scheduling and autoscaling for an internal AI platform usually comes down to designing around workload shape, queueing, and cluster economics. A good setup separates training, fine-tuning, and inference because they need very different scheduling policies.
1) Start by classifying workloads
Training / distributed jobs
- Long-running
- Need large contiguous GPU allocations
- Often tolerate queueing
- Benefit from gang scheduling and job priorities
Fine-tuning / batch jobs
- Medium duration
- Often bursty and interruptible
- Good candidates for preemption or spot instances
Inference / online serving
- Latency-sensitive
- Needs autoscaling based on QPS, latency, or GPU utilization
- Usually should be isolated from batch jobs
2) Use a scheduler that understands GPUs
If you’re on Kubernetes, common options are:
- Kubernetes + NVIDIA device plugin for basic GPU allocation
- Kueue for queueing and admission control for batch AI jobs
- Volcano for batch scheduling and gang scheduling
- Ray if your workloads are Python-native distributed jobs
- Slurm if you’re running more HPC-style training clusters
Recommended pattern
For most internal AI platforms:
- Kubernetes as the base
- Kueue or Volcano for training/fine-tuning jobs
- HPA/KEDA + custom metrics for inference autoscaling
- NVIDIA GPU Operator for driver/device management
3) Separate node pools by workload
Create distinct GPU node pools, for example:
- Inference pool
- Smaller GPUs or L4/T4/A10-class GPUs
- Higher availability
- Scale independently
- Training pool
- A100/H100 or equivalent
- Can scale aggressively
- Higher tolerance for queueing
- Spot/preemptible pool
- Best for non-urgent jobs
- Use checkpoints to handle eviction
Use:
- taints/tolerations
- node selectors
- affinity/anti-affinity
This prevents inference from being starved by training jobs.
4) Implement queueing and priority
GPU clusters fail operationally when jobs all try to start immediately.
Use:
- PriorityClasses for important workloads
- Quota / ResourceQuota per team or namespace
- Admission queues for fair sharing
- Preemption only when you truly need it
A common strategy:
- SRE/platform team gets reserved capacity
- Each product team gets quotas
- Lower-priority jobs can wait or run on spot nodes
5) Autoscaling for GPU nodes
There are two levels of autoscaling:
A. Pod autoscaling
For inference services:
- Scale pods based on:
- request rate
- p95 latency
- queue depth
- GPU utilization
- memory usage
Tools:
- HPA for CPU/memory/custom metrics
- KEDA for event-driven scaling
- Custom metrics adapter for Prometheus metrics
B. Node autoscaling
If pods can’t schedule because GPUs are full, scale nodes.
Tools:
- Cluster Autoscaler
- Karpenter on AWS
- Equivalent cloud autoscalers elsewhere
Important:
- GPU nodes often have long startup times
- Keep a small warm pool for latency-sensitive inference
- Use predictive scaling if workloads are periodic
6) Autoscaling strategy by workload
Inference
Scale on:
- QPS per replica
- p95/p99 latency
- queue depth
- GPU memory headroom
Best practices:
- Keep min replicas > 0
- Use canary deployments
- Prefer smaller model variants or quantized versions where possible
Batch training
Usually not pod-autoscaled; instead:
- Scale node pools based on queued jobs
- Use queue depth and pending GPU requests
- Batch systems should request whole GPUs explicitly
Multi-tenant fine-tuning
- Use quotas and max parallelism per team
- Add queue fairness
- Allow opportunistic burst beyond baseline capacity
7) Reserve and bin-pack intelligently
GPU scheduling can waste money if jobs are poorly packed.
Consider:
- Mixed GPU sizes only if your scheduler can handle fragmentation well
- Pack smaller inference workloads onto shared nodes when safe
- Use MIG on supported NVIDIA GPUs to split capacity for inference
- Avoid overcommitting GPU memory unless your framework supports it well
For NVIDIA A100/H100:
- MIG can help isolate inference workloads and improve utilization
8) Add observability before scaling problems hit
Track:
- GPU utilization
- GPU memory utilization
- Pending GPU jobs
- Queue wait time
- Node startup time
- Job success/failure rate
- Preemption/eviction count
- Cost per training run / inference request
Useful tooling:
- Prometheus + Grafana
- DCGM exporter for NVIDIA metrics
- Centralized logs for scheduling events
Alert on:
- pending jobs > threshold
- inference latency SLO breaches
- cluster autoscaler failing to provision nodes
- low GPU utilization despite high spend
9) Protect the platform from noisy neighbors
Use:
- namespace quotas
- per-team limits
- priority tiers
- separate inference and batch clusters if needed
For stricter isolation:
- dedicate certain node pools to certain teams or service classes
- apply network and IAM isolation as well
10) Practical reference architecture
A strong default design:
- Kubernetes
- NVIDIA GPU Operator
- Kueue or Volcano
- Cluster Autoscaler or Karpenter
- Prometheus + Grafana + DCGM exporter
- Separate GPU node pools:
- inference
- training
- spot/preemptible
- Priority classes, quotas, and queueing
- HPA/KEDA for inference services
11) Operational policies to define early
Write down policies for:
- who can request GPUs
- max GPUs per team
- priority order for workloads
- preemption rules
- checkpointing requirements for spot usage
- SLA/SLO expectations for inference
- when to reserve baseline capacity
12) A good default approach
If you want something pragmatic:
- Put GPUs on Kubernetes
- Split training and inference into separate node pools
- Use Kueue/Volcano for batch training queues
- Use HPA/KEDA for inference autoscaling
- Use Cluster Autoscaler/Karpenter for node scaling
- Add quotas, priorities, and observability
- Use spot GPUs for interruptible jobs
- Introduce MIG if inference utilization is poor
If you want, I can also give you:
- a reference Kubernetes architecture
- a sample autoscaling policy
- or a step-by-step rollout plan for an internal AI platform.