Prompt
How do I troubleshoot GPU jobs that fail because of driver mismatches?
Latest observation
GPU job failures from driver mismatches usually mean your code/container was built for a different NVIDIA driver/CUDA stack than what’s actually available on the machine or node.
Here’s a practical troubleshooting checklist.
1) Identify the exact error
Common symptoms:
CUDA driver version is insufficient for CUDA runtime versionfailed to initialize NVMLno kernel image is availableCUDA error: invalid device functioncould not select device driver "" with capabilities: [[gpu]]- container starts but sees no GPU
The exact message often tells you whether the issue is:
- host driver too old
- CUDA runtime too new
- container/runtime not wired to GPU
- wrong GPU architecture / build flags
2) Check the host driver version
On the machine or node:
nvidia-smi
Look at:
- Driver Version
- CUDA Version shown by
nvidia-smi(this is the max CUDA supported by the driver, not your installed toolkit)
Also useful:
cat /proc/driver/nvidia/version
3) Check what your job/container expects
Inside the environment where the job runs:
nvcc --version
python -c "import torch; print(torch.version.cuda)" # PyTorch
python -c "import tensorflow as tf; print(tf.sysconfig.get_build_info())" # TensorFlow
If you’re using a container, inspect its CUDA base image:
nvidia/cuda:12.x-runtimepytorch/pytorch:...-cuda12.x- etc.
Rule of thumb:
- Host driver must be new enough for the container/runtime CUDA version
- The NVIDIA driver comes from the host, not the container
4) Verify GPU access from the runtime
If using Docker:
docker run --rm --gpus all nvidia/cuda:12.4.1-base-ubuntu22.04 nvidia-smi
If that fails, the issue is usually with:
- NVIDIA Container Toolkit not installed/configured
- Docker not using the NVIDIA runtime
- permissions / cgroups / device plugin issues
Check:
nvidia-container-cli info
docker info | grep -i runtime
For Kubernetes:
kubectl describe pod <pod>
kubectl logs <pod>
kubectl describe node <node> | grep -A5 -i nvidia
Confirm:
- NVIDIA device plugin is running
- pod requests
nvidia.com/gpu - node has matching driver and toolkit setup
5) Compare CUDA/driver compatibility
If the driver is older than the CUDA runtime you need, you have two options:
- upgrade the host driver
- downgrade the container/runtime to a supported CUDA version
Example:
- Container uses CUDA 12.4
- Host driver only supports CUDA 12.1
- Fix by updating host driver or using a CUDA 12.1 image
6) Make sure the job isn’t mixing environments
Common in Python jobs:
- conda environment includes one CUDA version
- system libraries provide another
- pip wheel expects a different CUDA build
Check:
ldd your_binary | grep cuda
echo $LD_LIBRARY_PATH
python -c "import torch; print(torch.cuda.is_available())"
For PyTorch:
pip install torchwheels are usually bundled with a specific CUDA version- you generally do not need a local CUDA toolkit unless compiling extensions
7) Check GPU architecture compatibility
Sometimes the driver is fine, but the binary was compiled for the wrong GPU generation.
Symptoms:
invalid device functionno kernel image is available
Check GPU model:
nvidia-smi -L
Then ensure your build includes the right compute capability, especially if compiling custom CUDA kernels.
8) Review cluster scheduling issues
In shared clusters, a job can land on nodes with different driver versions.
Check node labels / partitions / queues:
- Are all GPU nodes on the same driver?
- Is your job pinned to a specific node type?
- Did the cluster recently upgrade some nodes but not others?
If so, use node selectors, partitions, or constraints to target compatible nodes.
9) Minimal reproducible test
Run a tiny test job on the same node/container:
nvidia-smi
python -c "import torch; print(torch.cuda.is_available()); print(torch.cuda.get_device_name(0))"
If this fails, the problem is environment/driver.
If it succeeds, the issue is likely in your application code or libraries.
10) If you manage the system
Recommended actions:
- Standardize GPU node driver versions
- Pin container images to known-compatible CUDA versions
- Use NVIDIA’s compatibility matrix
- Automate health checks on every GPU node:
nvidia-smi- container GPU smoke test
- driver/runtime version logging
Quick decision guide
- Host driver too old → upgrade driver or use older CUDA image
- Container can’t see GPU → fix NVIDIA runtime/toolkit configuration
- Compiled extension fails → rebuild for correct CUDA/GPU arch
- Only some nodes fail → cluster driver inconsistency or scheduling mismatch
If you want, I can also give you a step-by-step debugging flow for Docker, Kubernetes, Slurm, or PyTorch specifically.