Prompt
My CUDA versions keep breaking across containers; what should I check?
Latest observation
When CUDA “keeps breaking” across containers, it’s usually because one of a few version/compatibility boundaries is being crossed. Here’s what to check first.
1) Host driver vs container CUDA version
The host NVIDIA driver must be new enough for the CUDA runtime inside the container.
Check on the host:
nvidia-smi
Look at:
- Driver Version
- CUDA Version shown by
nvidia-smi(that’s the max CUDA level the driver supports)
Inside the container, check:
nvcc --version
# or
python -c "import torch; print(torch.version.cuda)"
Rule of thumb:
- Container CUDA can be older than the driver supports: usually fine.
- Container CUDA can be newer than what the driver supports: often breaks.
2) Are you accidentally mixing host CUDA libraries with container ones?
Containers should usually use the CUDA toolkit/runtime in the image, while relying on the host driver via NVIDIA Container Toolkit.
Check for accidental library overrides:
LD_LIBRARY_PATH- mounted
/usr/local/cuda - bind mounts from host into the container
- custom startup scripts
Inside the container:
echo $LD_LIBRARY_PATH
ls -l /usr/local/cuda
ldconfig -p | grep cuda
If host CUDA libraries are mounted in, they can conflict with container versions.
3) Base image tag and immutability
Using floating tags like:
latest12.4.0-runtimepytorch:2.4.0-cuda12.1-cudnn8without pinning digest
can lead to surprise changes.
Prefer:
- exact tags
- or even image digests
Example:
FROM nvidia/cuda:12.4.1-runtime-ubuntu22.04
Even better, pin by digest if reproducibility matters.
4) Framework CUDA version vs system CUDA version
If you use PyTorch, TensorFlow, JAX, etc., they often ship with their own CUDA/cuDNN expectations.
For PyTorch:
import torch
print(torch.version.cuda)
print(torch.cuda.is_available())
If PyTorch was built for CUDA 12.1 but your environment provides mismatched CUDA libraries, you can get runtime errors.
5) cuDNN / NCCL / libcublas version mismatches
CUDA issues are often actually caused by related libraries:
- cuDNN
- NCCL
- cuBLAS
- cuFFT
- driver/runtime compatibility
Symptoms:
- import errors
- “undefined symbol” errors
- NCCL hangs or collective failures
- runtime kernel launch failures
Check installed versions:
dpkg -l | egrep 'cuda|cudnn|nccl|cublas'
or in Python framework logs.
6) GPU architecture support
A container can have the right CUDA version but still fail if the binary doesn’t support your GPU’s compute capability.
Check GPU model:
nvidia-smi -L
Then confirm your framework/image supports that architecture.
7) NVIDIA Container Toolkit setup
Make sure the runtime is actually exposing the GPU correctly.
Test:
docker run --rm --gpus all nvidia/cuda:12.4.1-base-ubuntu22.04 nvidia-smi
If that fails, the issue is probably host/runtime configuration rather than CUDA itself.
Check:
nvidia-container-toolkitinstalled- Docker configured with NVIDIA runtime
--gpus allused- correct permissions on
/dev/nvidia*
8) Kernel modules / secure boot / host state
Sometimes the container is fine, but the host GPU stack is not:
- kernel update after driver install
- Secure Boot blocking modules
- reboot needed after driver changes
- stale driver modules loaded
Host checks:
nvidia-smi
lsmod | grep nvidia
9) Conda / pip library collisions
Inside containers, Python package installs can override or conflict with system CUDA libs.
Check for:
pip install nvidia-*- conda packages providing CUDA runtime
- mixed conda + system libraries
Use either:
- system CUDA runtime from image
- or conda-managed CUDA stack but avoid half-and-half unless you know the dependency graph.
10) Know the exact failure mode
Different errors point to different problems:
CUDA driver version is insufficient for CUDA runtime version→ driver too old for container CUDAlibcudart.so not found→ missing runtime library path/image issueundefined symbol: ...→ library mismatchno kernel image is available for execution→ GPU architecture mismatchNCCL error→ networking/NCCL/version issue
Quick checklist
Run these on host and in container:
Host
nvidia-smi
Container
nvidia-smi
nvcc --version
echo $LD_LIBRARY_PATH
python -c "import torch; print(torch.version.cuda, torch.cuda.is_available())"
Then verify:
- driver supports container CUDA
- no host CUDA libs are leaking in
- image tags are pinned
- cuDNN/NCCL versions are compatible
- NVIDIA Container Toolkit is correctly configured
If you want, paste:
- host
nvidia-smi - container
nvcc --versionortorch.version.cuda - the exact error
and I can help pinpoint the mismatch.