Prompt

How do I monitor GPU utilization and queue time?

Artificial Intelligence / AI Infrastructure1 observationLast seen Oct 11, 2026

Latest observation

Oct 11, 2026OpenAI APIWeb search: off

To monitor GPU utilization and queue time, you typically track two different things:

1) GPU utilization

This tells you how busy the GPU is while jobs are running.

Common tools

  • nvidia-smi (NVIDIA GPUs)

    nvidia-smi
    

    Shows:

    • GPU memory used
    • GPU utilization %
    • active processes
  • Continuous monitoring

    watch -n 1 nvidia-smi
    
  • More detailed metrics

    nvidia-smi dmon
    

    or

    nvidia-smi pmon
    

    These give per-second/per-process performance stats.

In code

If you’re using Python, libraries like:

  • pynvml
  • torch.cuda (limited, more for availability/memory than full utilization)

Example with pynvml:

from pynvml import *
nvmlInit()
handle = nvmlDeviceGetHandleByIndex(0)
print(nvmlDeviceGetUtilizationRates(handle).gpu)

2) Queue time

This is the time a job waits before it starts running.

If you’re using a scheduler

  • SLURM: queue time = start_time - submit_time
    • You can inspect jobs with:
      squeue -u $USER
      scontrol show job <jobid>
      sacct -j <jobid> --format=JobID,Submit,Start,Elapsed
      
  • Kubernetes: queue time = time from pod creation to container start
    • Inspect:
      kubectl describe pod <pod>
      kubectl get pod <pod> -o json
      

If you manage your own system

You’ll need to log:

  • submission timestamp
  • actual start timestamp

Then compute:

queue_time = start_time - submit_time

3) Best practice: monitor both together

For most setups, use:

  • GPU monitoring: nvidia-smi, Prometheus + Grafana, or cloud monitoring
  • Queue monitoring: scheduler logs / job metadata

If you’re on a cluster, a common setup is:

  • Prometheus scraping GPU exporters like dcgm-exporter
  • Grafana dashboards for GPU load, memory, temperature, and job wait times

4) If you want a dashboard

Useful stack:

  • NVIDIA DCGM Exporter
  • Prometheus
  • Grafana

This gives:

  • GPU utilization
  • memory usage
  • power/temperature
  • per-node status
  • historical trends

Queue time usually comes from your scheduler or application logs rather than GPU telemetry.


If you tell me your environment (e.g. SLURM, Kubernetes, AWS, or a single machine), I can give you exact commands or a ready-to-use monitoring setup.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.