Prompt
How do I manage CUDA and driver versions across GPU nodes?
Latest observation
Managing CUDA and NVIDIA driver versions across GPU nodes is mostly about decoupling what you can, standardizing what you can’t, and automating validation.
Key rule: driver vs CUDA toolkit
- NVIDIA driver runs on the host and is required for the GPU to work.
- CUDA toolkit (nvcc, libraries, headers) is usually for building/running CUDA apps and can often be managed separately.
- In many setups, the driver must be new enough for the CUDA runtime/toolkit you use, but the toolkit does not need to exactly match the driver.
Recommended strategy
1) Pick a supported compatibility target
Define one or more approved combinations:
- OS version
- NVIDIA driver version
- CUDA toolkit version
- cuDNN / NCCL versions if relevant
Use NVIDIA’s compatibility matrix to ensure:
- the driver supports the desired CUDA version
- your frameworks (PyTorch, TensorFlow, RAPIDS, etc.) support that CUDA version
2) Standardize the host driver
Treat the driver as part of the node base image or node bootstrap process:
- Use the same driver branch across all nodes in a cluster when possible
- Install via:
- OS package manager (preferred for manageability), or
- vendor runfile only if required
- Pin versions so updates are deliberate, not accidental
Good practices:
- keep driver versions consistent within a cluster
- update in a controlled rollout
- reboot and verify after upgrades
3) Use containers for CUDA user-space
A very common pattern:
- Host: NVIDIA driver only
- Container: CUDA toolkit + app dependencies
This works well with:
- NVIDIA Container Toolkit
- CUDA base images like
nvidia/cuda:XX.YY-runtimeordevel
Benefits:
- lets different jobs use different CUDA versions on the same host
- reduces node drift
- makes reproducibility easier
4) Avoid installing multiple CUDA toolkits system-wide unless necessary
If you need multiple CUDA versions:
- prefer containers, or
- install side-by-side under separate prefixes like:
/usr/local/cuda-11.8/usr/local/cuda-12.4
- use environment modules (
module load cuda/12.4) to select versions
5) Keep applications pinned to known-good stacks
For each workload, define:
- driver minimum
- CUDA runtime/toolkit version
- framework version
- cuDNN/NCCL version
Examples:
- PyTorch wheel built for CUDA 12.1
- TensorFlow build requiring CUDA 12.x and specific cuDNN
- HPC application built against CUDA 11.8
6) Validate on every node
Add a health check after provisioning and before scheduling workloads:
nvidia-sminvcc --versionif toolkit installed- small CUDA sample or framework smoke test
- confirm driver/runtime compatibility
Example checks:
nvidia-smi
cat /proc/driver/nvidia/version
nvcc --version
7) Automate with configuration management
Use tools like:
- Ansible
- Puppet
- Chef
- Salt
- cloud-init
- image baking (Packer)
This helps you:
- install exact versions
- pin packages
- validate post-install
- roll back if needed
8) Plan upgrades carefully
When upgrading:
- upgrade in staging first
- verify apps/containers work
- roll out in batches
- keep a rollback path
- watch for kernel/driver module compatibility after OS updates
Common deployment patterns
Pattern A: Bare metal with host driver + containers
Best for most modern AI/HPC clusters.
- Node image includes driver
- Workloads run in containers with CUDA toolkit
- Different containers can use different CUDA versions
Pattern B: Shared system CUDA install
Useful for smaller clusters or legacy environments.
- Install one or more CUDA versions on the host
- Manage via modules/environment variables
- More operational overhead
Pattern C: Immutable node images
Best for large fleets.
- Bake OS + driver into a golden image
- Rebuild images for upgrades
- Reduces configuration drift
Practical tips
- Don’t assume
latestis safe. - Don’t mix random toolkit versions with an old driver.
- Use the minimum driver version required by your chosen CUDA release.
- For containers, make sure the host has:
- NVIDIA driver
- NVIDIA Container Toolkit
- On multi-user systems, environment modules are often easier than manual
PATHedits.
Suggested operational policy
A simple policy that works well:
- One approved driver version per cluster generation
- One or two supported CUDA versions
- Containerize application dependencies
- Bake drivers into images or install via automation
- Test on staging before production rollout
If you want, I can also give you:
- a version compatibility cheat sheet,
- an Ansible example for driver/CUDA management, or
- a recommended setup for Kubernetes / Slurm / bare metal.