Prompt
Why are my GPU nodes idle but inference still queues up?
Latest observation
Oct 9, 2026
OpenAI APIWeb search: off
This usually means the GPU nodes are “up,” but not actually available to the scheduler or serving layer. Common causes:
1) Pods are not being scheduled onto the GPU nodes
Even if the nodes are idle, inference pods may be:
- pending due to insufficient GPU requests/limits
- blocked by taints/tolerations
- constrained by node selectors / affinity
- stuck because the GPU device plugin isn’t advertising GPUs
Check:
kubectl get pods -A -o widekubectl describe pod <pending-pod>kubectl describe node <gpu-node> | grep -A5 Allocatable
2) The service has no healthy ready replicas
Requests can queue if:
- inference workers are starting slowly
- readiness probes fail
- pods crashloop
- autoscaler hasn’t added enough replicas yet
Check:
kubectl get deploy,podskubectl describe pod <inference-pod>- logs for model load errors or OOMs
3) You have a queue in the inference server itself
Some servers intentionally queue requests when:
- max concurrent requests is low
- batcher is configured aggressively
- worker pool is smaller than request rate
- model instances are pinned to a single GPU
Examples:
- Triton model instance count too low
- vLLM
max_num_seqs/max_num_batched_tokenslimits - TGI limited by
--max-concurrent-requests
4) GPU utilization is low, but memory or another bottleneck is saturated
A GPU can appear “idle” by compute metrics while actually blocked by:
- GPU memory fragmentation / OOM pressure
- CPU tokenization bottleneck
- network or disk I/O
- single-threaded pre/post-processing
- host RAM or PCIe contention
Check:
nvidia-smi- CPU usage on the inference pods
- model server logs and latency breakdowns
5) Autoscaling or load balancing is misconfigured
Requests may queue if:
- traffic is sent to one pod instead of spread across replicas
- HPA uses CPU metrics only, not request queue length
- the service mesh / ingress has limited concurrency
- the load balancer still routes to terminating/not-ready pods
6) The GPU nodes are idle because they’re reserved or unreachable
Nodes may look idle but be unusable if:
NotReady- cordoned/drained
- lacking drivers/CUDA/runtime
- missing the correct MIG configuration
- device plugin daemonset not running
Fastest way to narrow it down
- Are inference pods Pending or Running?
- Are the pods Ready?
- Do the pods request GPUs?
- Does
nvidia-smishow the process on the GPU? - Is the queue in the app/server or in Kubernetes scheduling?
If you want, paste:
kubectl get pods -A -o widekubectl describe pod <one-queued-inference-pod>kubectl describe node <gpu-node>- your inference server type (Triton/vLLM/TGI/custom)
and I can help pinpoint the exact bottleneck.