Prompt
How do I figure out why my GPU nodes are idling?
Latest observation
To figure out why GPU nodes are idling, you want to determine whether the problem is:
- No work is being requested
- Work is requested but can’t be scheduled
- Work is scheduled but not using the GPU
- The GPU is being underutilized by the workload itself
Here’s a practical debugging checklist.
1) Confirm the GPUs are actually idle
On the node:
nvidia-smi
Check:
- GPU utilization %
- Memory usage
- Running processes
If utilization is near 0 and no processes are listed, the GPU is truly idle.
For more detail:
watch -n 1 nvidia-smi
2) See whether the cluster has pending GPU work
If you’re using Kubernetes:
kubectl get pods -A -o wide
kubectl describe pod <pod>
kubectl get events -A --sort-by=.lastTimestamp
Look for:
- Pods stuck in
Pending - Scheduling errors like:
Insufficient nvidia.com/gpu- node selector / affinity mismatch
- taints/tolerations issues
If you use a batch system like Slurm:
squeue
scontrol show job <jobid>
3) Check if the nodes are advertising GPUs correctly
For Kubernetes:
kubectl describe node <node>
Look under allocatable/resources for GPU count, e.g.:
nvidia.com/gpu: 4
If the GPU resource is missing, the device plugin may be broken.
Also check:
kubectl get daemonset -A | grep -i nvidia
kubectl logs -n kube-system -l name=nvidia-device-plugin-ds
4) Check whether the scheduler can place GPU workloads
Common blockers:
- Node taints without tolerations
- Wrong node labels / selectors
- Pod requests GPU but also requests incompatible CPU/memory
- Too restrictive affinity/anti-affinity
- Namespace quota limits
Example:
kubectl describe pod <pod>
Look for the scheduler’s exact failure reason.
5) Make sure the workload is actually using the GPU
Sometimes a job lands on a GPU node but runs on CPU only.
Check inside the container:
nvidia-smi
And for frameworks:
- PyTorch:
torch.cuda.is_available() torch.cuda.device_count() - TensorFlow:
tf.config.list_physical_devices('GPU')
Common issues:
- Missing CUDA libraries
- Wrong image
- Model/device not moved to GPU
- Batch size too small
- Data loader bottlenecking CPU/IO
6) Look for resource fragmentation
You may have GPUs free, but not enough free memory or the workload requests too much of a certain node type.
For Kubernetes, check:
kubectl describe node <node>
kubectl top nodes
kubectl top pods -A
A GPU node can appear idle if:
- CPU is exhausted
- memory is fragmented
- pods are pinned to specific GPU types
7) Verify autoscaling or queueing behavior
If you rely on autoscaling:
- Is the cluster autoscaler enabled?
- Is it allowed to scale GPU node groups?
- Are pending GPU pods triggering scale-up?
Check autoscaler logs and node group configuration. Sometimes the issue is that pending pods never trigger scale-up because of:
- unschedulable pod constraints
- node group min/max limits
- quota exhaustion
- missing GPU instance capacity in the cloud region
8) Inspect workload-level bottlenecks
A GPU can idle if the job is waiting on:
- disk / network I/O
- CPU preprocessing
- synchronization overhead
- too small microbatches
- inefficient kernels
Look at:
- GPU utilization vs memory usage
- CPU usage on the same pod
- input pipeline latency
- profiler traces
Useful tools:
nvidia-smi dmon- Nsight Systems / Nsight Compute
- framework profilers
9) Check for admission / policy / quota issues
In Kubernetes:
kubectl get resourcequota -A
kubectl get limitrange -A
Also review:
- Pod Security policies / admission webhooks
- image pull failures
- init container failures
- missing secrets/configmaps
10) Follow the signal path
A good way to isolate it is:
Is there a job waiting?
- No → nothing to run
- Yes → continue
Can it schedule?
- No → scheduling/constraint issue
- Yes → continue
Does it start on GPU?
- No → image/runtime/device plugin issue
- Yes → continue
Does it use the GPU?
- No → app/config/code issue
- Yes but low utilization → workload bottleneck
Quick Kubernetes commands to start with
kubectl get pods -A | grep Pending
kubectl describe pod <pending-pod>
kubectl describe node <gpu-node>
kubectl logs -n kube-system -l name=nvidia-device-plugin-ds
kubectl get events -A --sort-by=.lastTimestamp
If you tell me what environment you’re using — Kubernetes, Slurm, Ray, SageMaker, Vertex, or plain VMs — I can give you a targeted troubleshooting flow and the exact commands to run.