Prompt
Why do my containers fail to see the GPU on this server?
Latest observation
Oct 11, 2026
OpenAI APIWeb search: off
Most often this happens because the GPU is visible to the host, but the container runtime isn’t actually passing the device into the container.
Common causes:
-
No NVIDIA container runtime / toolkit installed
- Docker/Podman won’t expose GPUs unless NVIDIA Container Toolkit is installed and configured.
-
Container started without GPU access
- For Docker you usually need something like:
docker run --gpus all ... - Without
--gpus, the container won’t see the GPU.
- For Docker you usually need something like:
-
Wrong container image
- The image may not include NVIDIA user-space libraries or CUDA support.
- Some images are CPU-only even if run on a GPU host.
-
Driver mismatch
- The host has a driver, but the container expects different CUDA/runtime versions.
- Usually the host driver must be new enough for the CUDA version in the container.
-
Permissions / cgroup / device access issue
- The container may not have access to
/dev/nvidia*devices. - Security settings, rootless Docker, or restrictive cgroups can block them.
- The container may not have access to
-
GPU driver not healthy on the host
- Check the host first:
nvidia-smi - If that fails on the host, containers won’t work either.
- Check the host first:
Quick checks:
- On host:
nvidia-smi docker info | grep -i nvidia - In container:
ls -l /dev/nvidia* nvidia-smi
If you want, I can help you debug it step-by-step if you paste:
- host OS
- container runtime (Docker/Podman/Kubernetes)
- output of
nvidia-smion the host - the exact container run command or compose file