Prompt
How do I monitor GPU utilization and queue time?
Latest observation
Oct 11, 2026
OpenAI APIWeb search: off
To monitor GPU utilization and queue time, you typically track two different things:
1) GPU utilization
This tells you how busy the GPU is while jobs are running.
Common tools
-
nvidia-smi(NVIDIA GPUs)nvidia-smiShows:
- GPU memory used
- GPU utilization %
- active processes
-
Continuous monitoring
watch -n 1 nvidia-smi -
More detailed metrics
nvidia-smi dmonor
nvidia-smi pmonThese give per-second/per-process performance stats.
In code
If you’re using Python, libraries like:
pynvmltorch.cuda(limited, more for availability/memory than full utilization)
Example with pynvml:
from pynvml import *
nvmlInit()
handle = nvmlDeviceGetHandleByIndex(0)
print(nvmlDeviceGetUtilizationRates(handle).gpu)
2) Queue time
This is the time a job waits before it starts running.
If you’re using a scheduler
- SLURM: queue time =
start_time - submit_time- You can inspect jobs with:
squeue -u $USER scontrol show job <jobid> sacct -j <jobid> --format=JobID,Submit,Start,Elapsed
- You can inspect jobs with:
- Kubernetes: queue time = time from pod creation to container start
- Inspect:
kubectl describe pod <pod> kubectl get pod <pod> -o json
- Inspect:
If you manage your own system
You’ll need to log:
- submission timestamp
- actual start timestamp
Then compute:
queue_time = start_time - submit_time
3) Best practice: monitor both together
For most setups, use:
- GPU monitoring:
nvidia-smi, Prometheus + Grafana, or cloud monitoring - Queue monitoring: scheduler logs / job metadata
If you’re on a cluster, a common setup is:
- Prometheus scraping GPU exporters like
dcgm-exporter - Grafana dashboards for GPU load, memory, temperature, and job wait times
4) If you want a dashboard
Useful stack:
- NVIDIA DCGM Exporter
- Prometheus
- Grafana
This gives:
- GPU utilization
- memory usage
- power/temperature
- per-node status
- historical trends
Queue time usually comes from your scheduler or application logs rather than GPU telemetry.
If you tell me your environment (e.g. SLURM, Kubernetes, AWS, or a single machine), I can give you exact commands or a ready-to-use monitoring setup.