Prompt
Why do my GPU containers keep crashing after deploys?
Latest observation
GPU containers that crash right after deploys are usually failing for one of a few common reasons. The key is to determine whether the crash is happening because of GPU access, driver/runtime mismatch, resource pressure, or the app itself.
Common causes
1. NVIDIA driver / CUDA mismatch
If the host driver, container CUDA version, and app expectations don’t line up, the container may start and then crash when it touches the GPU.
Typical symptoms:
- Works on one node but not another
nvidia-smifails inside the container- Logs mention
CUDA error,cuInit failed,libcuda.so, ordriver version is insufficient
What to check:
- Host NVIDIA driver version
- Container base image CUDA version
- Whether the node actually has the NVIDIA container runtime configured
2. GPU device not exposed to the container
Your deploy may be landing on a node without GPUs, or the pod/container may not be requesting them correctly.
Typical symptoms:
- Container runs but falls back badly or exits
- Kubernetes pod stays in
CrashLoopBackOfforPending - Logs show no visible CUDA devices
What to check:
- Kubernetes GPU requests, e.g.
nvidia.com/gpu: 1 - Node selectors / tolerations / affinity
- Whether the deployment is scheduled onto GPU-capable nodes only
3. Container starts before GPU drivers are ready
After a deploy or node restart, the container may come up before the GPU stack is fully initialized.
Typical symptoms:
- Intermittent failures after rollout
- First few pods fail, later ones work
- Works after retry or manual restart
What to check:
- Node reboot timing
- NVIDIA device plugin status
- Readiness/liveness probe timing
4. Memory exhaustion: GPU or CPU
A model load can fail if GPU memory is too tight, or the container can be OOM-killed by the host.
Typical symptoms:
- Sudden exit with no clear app error
OOMKilledin KubernetesCUDA out of memory- Crashes only after deploy when new version uses more memory
What to check:
kubectl describe podforOOMKilled- GPU memory usage via
nvidia-smi - CPU/RAM limits vs actual usage
5. Health checks are too aggressive
A deploy may pass traffic before the model is loaded, then probes kill it.
Typical symptoms:
- Container starts, then restarts repeatedly
- Liveness probe fails during model warmup
- Crashes only during rollout
What to check:
- Startup time of the app
- Liveness/readiness probe thresholds
- Add a
startupProbeif on Kubernetes
6. Application initialization errors
The new deploy may have a code/config issue unrelated to GPU, but it appears as a GPU crash because the service only fails when initializing models.
Typical symptoms:
- Errors loading weights, configs, secrets, or paths
- Dependency mismatch in Python packages
- Different behavior between old and new image tags
What to check:
- Container logs from startup
- Diff between image versions
- Missing files, environment variables, or secret mounts
7. Node-level GPU instability or driver reset
Sometimes the node’s GPU driver resets or the GPU gets into a bad state after deploys or load spikes.
Typical symptoms:
- Multiple pods on same node crash
- Kernel logs show NVIDIA/Xid errors
dmesgor node logs mention GPU reset
What to check:
- Node logs
nvidia-smi -q- Whether the problem follows the node rather than the image
Fast debugging checklist
-
Look at pod status
kubectl get podskubectl describe pod <pod>- Check for
OOMKilled, probe failures, or scheduling issues
-
Inspect container logs
kubectl logs <pod> --previous- Look for CUDA, driver, or model-loading errors
-
Verify GPU visibility
- Exec into the pod and run
nvidia-smi - Confirm the container can see the GPU
- Exec into the pod and run
-
Check node and driver versions
- Host driver version
- Image CUDA version
- NVIDIA device plugin / runtime status
-
Check probes and startup timing
- Temporarily relax liveness/readiness probes
- Add a startup probe if needed
-
Check scheduling
- Ensure the pod lands on GPU nodes only
- Confirm GPU resource requests are set
Very common Kubernetes-specific fix
If this is on Kubernetes, make sure you have something like:
resources:
limits:
nvidia.com/gpu: 1
And ensure the node has:
- NVIDIA driver installed
- NVIDIA Container Toolkit
- NVIDIA device plugin
- Correct runtime class or default GPU runtime, depending on setup
If you want the quickest root-cause path
Send me:
- The pod
describeoutput - The last 50 lines of container logs
- Your deployment YAML resource/probe section
- Host GPU driver version and container CUDA version
With that, I can usually tell you the likely cause very quickly.
Brands