Prompt
Why do model deployments on Kubernetes keep crashing?
Latest observation
Kubernetes model deployments usually keep crashing for one of a few common reasons. The exact cause depends on whether the pod is CrashLoopBackOff, OOMKilled, stuck in ImagePullBackOff, failing readiness/liveness probes, or exiting immediately.
Most common causes
1. Out of memory
Your model container may be using more RAM than the pod limit.
- Symptoms:
OOMKilled, restarts after loading the model - Common with large LLMs, big batches, or Python workers with high overhead
Fixes
- Increase memory limits/requests
- Reduce model size, batch size, or concurrency
- Use quantization or smaller model variants
- Make sure the app doesn’t preload multiple copies of the model
2. CPU throttling or startup timeout
The pod may be too slow to start because the model takes a long time to load.
- Symptoms: liveness probe fails before model is ready
- Kubernetes kills the container repeatedly
Fixes
- Add a longer
startupProbe - Increase
initialDelaySecondsfor probes - Make liveness probe less aggressive
- Optimize model loading or warm start
3. Bad health probes
If your app returns 500/timeout on /health until the model is loaded, probes can kill it.
- Symptoms: container starts, then restarts every few seconds
Fixes
- Use a
startupProbefor slow initialization - Ensure readiness probe checks only “ready to serve,” not “fully warmed up”
- Ensure liveness probe checks only process health, not model availability
4. Missing environment variables or config
The app may crash on startup because required values aren’t set.
- Examples:
MODEL_PATH, API keys, database URLs, bucket names
Fixes
- Check container logs
- Verify ConfigMaps/Secrets are mounted and referenced correctly
- Validate startup scripts locally with the same env vars
5. Wrong container command/entrypoint
The image may be fine, but Kubernetes is launching the wrong command.
- Symptoms: container exits immediately with code 1 or 127
Fixes
- Confirm
commandandargsin the manifest - Check if the image expects
python app.py,uvicorn ..., or a shell wrapper - Make sure executable paths exist in the image
6. Dependency/import failures
Python apps often fail because a required library is missing or incompatible.
- Symptoms: traceback in logs, crash on import
Fixes
- Check logs with
kubectl logs <pod> - Ensure the image was built with the correct dependencies
- Pin versions for model frameworks like PyTorch, TensorFlow, CUDA, transformers
7. GPU/runtime mismatch
If the model needs GPU, but the node or image doesn’t match, it can crash.
- Symptoms: CUDA errors,
nvidia-smiunavailable, device not found
Fixes
- Confirm GPU nodes exist and the pod requests a GPU
- Ensure the container image matches the CUDA runtime on the node
- Verify NVIDIA device plugin is installed
8. Volume/mount issues
Model files may not be present or accessible.
- Symptoms: file not found, permission denied, empty mount
Fixes
- Check PVC status and mount paths
- Verify permissions and filesystem ownership
- Ensure the model artifact is downloaded before startup if required
9. Container exits because the app is designed as a batch job
Sometimes the container finishes work and exits, but Kubernetes expects a long-running service.
Fixes
- If it’s a server, make sure the process stays alive
- If it’s batch inference, use a
Jobinstead of aDeployment
10. Node pressure or eviction
The pod may be evicted due to node memory/disk pressure.
- Symptoms: pod disappears or restarts on a different node
Fixes
- Check node conditions
- Reduce resource usage
- Add resource requests/limits that match reality
What to check first
Run these:
kubectl get pods
kubectl describe pod <pod-name>
kubectl logs <pod-name> --previous
Look for:
Last StateReason: OOMKilled- probe failures
- exit code
- import errors / stack traces
Also check events:
kubectl get events --sort-by=.metadata.creationTimestamp
A practical rule of thumb
For model serving pods, crashing is often caused by:
- not enough memory
- startup probes too strict
- missing config/dependencies
- GPU/runtime mismatch
If you want, paste:
kubectl describe pod ...- the last 50 lines of
kubectl logs ... --previous - your deployment YAML
and I can help pinpoint the exact reason.
Brands