Prompt
How do I debug GPU memory errors during inference deployment?
Latest observation
Here’s a practical way to debug GPU memory errors during inference deployment.
1) First identify the exact failure type
Common GPU memory-related errors include:
- Out of memory (OOM): the model or batch doesn’t fit in VRAM.
- CUDA illegal memory access: often a kernel bug, shape mismatch, or a custom op issue.
- Memory leak / gradual growth: memory increases after each request or batch.
- Fragmentation: enough total free memory exists, but not in a contiguous block.
Check the full error message and stack trace. The wording matters a lot.
2) Measure baseline memory usage
Before changing code, measure:
- GPU total memory
- memory used by:
- model weights
- activations / intermediates
- KV cache for LLMs
- framework overhead
- request batching overhead
Useful tools:
nvidia-smi- PyTorch:
torch.cuda.memory_allocated()torch.cuda.memory_reserved()torch.cuda.max_memory_allocated()torch.cuda.memory_summary()
Example:
import torch
print(torch.cuda.memory_allocated() / 1024**2, "MB allocated")
print(torch.cuda.memory_reserved() / 1024**2, "MB reserved")
print(torch.cuda.max_memory_allocated() / 1024**2, "MB peak allocated")
print(torch.cuda.memory_summary())
If allocated is low but reserved is high, fragmentation or allocator behavior may be involved.
3) Reproduce with a minimal case
Try to isolate:
- one model
- one request
- one batch size
- fixed input shape/length
If the error disappears with a smaller batch or shorter sequence, it’s likely a true capacity issue.
For LLMs, test separately:
- prompt length
- generation length
- number of concurrent requests
- max batch tokens
4) Check the obvious capacity knobs
Depending on your stack, reduce one or more of:
- batch size
- max sequence length / context length
- number of concurrent in-flight requests
- tensor parallel degree
- beam size
- image resolution / input resolution
- number of sampled outputs
For LLM serving, the KV cache is often the real memory driver, not just the model weights.
5) Determine whether the issue is peak memory or a leak
Peak memory issue
The app fails only on large inputs or high concurrency.
Fixes:
- reduce batch size
- use mixed precision (
fp16/bf16) - quantize model weights
- offload some layers / KV cache
- use chunked prefill or paged attention
- decrease max tokens per request
Leak issue
Memory grows over time across requests.
Common causes:
- tensors kept in global lists/caches
- storing outputs without
.detach()or.cpu() - accumulating computation graphs during inference
- not using
torch.no_grad()ortorch.inference_mode() - framework/model server bug
- per-request objects not freed
Use a loop and monitor memory after each request.
Example:
with torch.inference_mode():
out = model(inp)
6) Look for hidden tensor retention
A very common bug is accidentally keeping GPU tensors alive.
Check for:
- appending outputs to a Python list
- logging tensors without converting to CPU/NumPy
- closures or caches holding references
- storing hidden states, logits, or attention maps
- returning tensors that remain referenced by the caller
If you need to keep results, move them off GPU:
result = output.detach().cpu()
7) Verify inference mode is actually inference mode
For PyTorch:
- use
model.eval() - use
torch.inference_mode()or at leasttorch.no_grad()
model.eval()
with torch.inference_mode():
y = model(x)
Without this, autograd may track graphs and increase memory significantly.
8) Inspect allocation patterns and fragmentation
If you see OOM with free memory still available:
- PyTorch allocator fragmentation may be involved
- variable input sizes can fragment memory
- long-running services can accumulate fragmented blocks
Things to try:
- use more consistent input shapes
- enable memory-efficient batching
- restart workers periodically
- adjust allocator settings, if appropriate
- in PyTorch, experiment with:
PYTORCH_CUDA_ALLOC_CONF=max_split_size_mb:128expandable_segments:Truein supported versions
9) Check for invalid memory access in custom CUDA ops
If the error is illegal memory access, device-side assert, or the process crashes after a kernel error:
- run with:
CUDA_LAUNCH_BLOCKING=1
- check tensor shapes, dtypes, and index ranges
- validate custom kernels and extension code
- confirm model inputs match expected ranges
- look for out-of-bounds indexing in embeddings or token IDs
This class of bug is often not “memory pressure” but a correctness issue.
10) Use profiling tools
If the issue is unclear, profile:
- Nsight Systems for timeline and concurrency
- Nsight Compute for kernel-level details
- PyTorch profiler for memory and operator-level analysis
This helps identify:
- which op spikes memory
- which request shape triggers the spike
- whether memory is in model forward, pre/post-processing, or caching
11) Consider model-serving-specific causes
For deployment systems like Triton, vLLM, TensorRT-LLM, Hugging Face TGI, etc., memory issues often come from:
- too-large prefill/prompt batches
- KV cache sizing
- continuous batching policy
- multiple model replicas on one GPU
- warmup requests using large shapes
- CPU-GPU transfer buffers
- tokenizer or preprocessor memory not released
Check server configs for:
- max batch tokens
- max input length
- max output length
- number of concurrent sequences
- GPU memory utilization target
- swap/offload settings
12) A simple debugging checklist
Use this order:
- Reproduce the failure with one request.
- Reduce batch size and sequence length.
- Confirm
eval()+inference_mode(). - Print
allocated/reserved/max_memory_*. - Check for stored GPU tensors.
- Profile the failing request.
- Test for fragmentation with more consistent shapes.
- Validate any custom CUDA ops.
- Review serving config for KV cache and concurrency limits.
13) Quick mitigations if you need to stabilize production now
- lower batch size or concurrency
- shorten max input/output lengths
- switch to
fp16orbf16 - quantize weights
- enable paged attention / KV cache optimization
- use smaller model or more GPUs
- restart workers periodically if there’s a leak you haven’t fixed yet
- add request-level guards to reject oversized inputs
If you want, I can also give you:
- a PyTorch-specific GPU memory debugging script,
- a checklist for LLM serving (vLLM/TGI/TensorRT-LLM), or
- help interpret a specific OOM / illegal memory access error message.