Prompt
How do I run inference on GPUs with predictable latency?
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
To get predictable latency for GPU inference, the main idea is to reduce sources of variability: dynamic work, kernel launches, memory transfers, CPU overhead, and contention. Here’s a practical checklist.
1) Keep the model and inputs fixed
- Use a fixed batch size if possible, ideally batch size 1 for latency-sensitive serving.
- Avoid dynamic input shapes; use static shapes or a small set of precompiled shapes.
- Keep sequence lengths capped and padded consistently for LLMs.
2) Warm up everything
Before measuring or serving:
- Run several warmup inferences to:
- trigger CUDA context initialization
- load kernels into cache
- populate autotuned algorithm choices
- compile any graphs / JIT paths
- If using TensorRT / Torch / XLA / TVM, build or cache the optimized engine ahead of time.
3) Use an optimized inference runtime
Preferred options:
- TensorRT for NVIDIA GPUs and static or semi-static models
- ONNX Runtime + CUDA/TensorRT Execution Provider
- PyTorch 2 compile / TorchScript where applicable
- vLLM / TensorRT-LLM / FasterTransformer for LLM serving
These reduce overhead and often fuse ops, improving consistency.
4) Preallocate memory
- Preallocate input/output buffers on GPU.
- Reuse buffers across requests.
- Avoid per-request
cudaMalloc/free. - Use pinned host memory for CPU↔GPU transfers.
- If possible, keep preprocessing on GPU too.
5) Minimize host-device transfers
Transfers add latency and jitter.
- Send data to GPU once, keep it there if you can.
- Batch CPU preprocessing and use pinned memory.
- Overlap copies with compute only if it doesn’t complicate tail latency.
- Prefer an all-GPU pipeline for decoding/postprocessing.
6) Avoid CPU-side bottlenecks
Latency variance often comes from the CPU, not the GPU.
- Pin inference worker threads to specific CPU cores.
- Use a dedicated inference process.
- Avoid noisy neighbors and background tasks.
- Use a separate CPU thread pool for preprocessing if needed.
- Don’t let Python GIL-heavy code sit on the critical path.
7) Control GPU contention
Predictability needs isolation:
- Reserve a GPU for the service if possible.
- Avoid sharing with training jobs or other inference workloads.
- If multi-tenant, use MIG on supported NVIDIA GPUs for hardware isolation.
- Avoid oversubscribing CUDA streams unless you’ve profiled it carefully.
8) Use CUDA graphs if the workload is repetitive
For fixed-shape, repeated inference:
- CUDA Graphs can reduce launch overhead and jitter.
- Very useful for small models or low-latency requests.
- Requires stable tensor shapes and memory addresses.
9) Disable variable behavior where possible
For consistent latency:
- Turn off unnecessary debug/logging.
- Fix precision mode: FP16, BF16, or INT8 with calibrated scales.
- Avoid runtime autotuning during serving.
- Use deterministic modes only if they don’t hurt performance too much.
10) Measure p95/p99, not just average
Predictable latency means looking at tail latency:
- Track p50, p95, p99, p99.9
- Measure under realistic load
- Include:
- queueing delay
- preprocessing
- GPU compute
- postprocessing
- network overhead
11) Consider request scheduling carefully
- For strict latency, prefer no batching or very small micro-batches.
- If batching is needed, use bounded dynamic batching with a very short max wait time.
- Separate interactive and bulk traffic.
12) Power and thermal stability matter
GPU clock throttling creates jitter.
- Ensure adequate cooling and power headroom
- Consider locking application clocks where supported
- Avoid thermal throttling by monitoring temperature and power draw
A good low-latency setup looks like this
- Static model / fixed input shapes
- TensorRT or similar optimized engine
- Preallocated buffers
- Pinned CPU threads
- Warmed-up engine
- Dedicated GPU or MIG slice
- CUDA graphs if applicable
- No or minimal batching
- Continuous p99 monitoring
Example pattern
If you’re serving a CNN or transformer:
- Precompile the model to TensorRT.
- Load the engine at startup.
- Allocate all buffers once.
- Warm up with 10–100 dummy requests.
- Serve with fixed shapes and a dedicated GPU.
- Monitor tail latency and GPU utilization.
If you want, I can give you:
- a PyTorch low-latency inference template
- a TensorRT deployment checklist
- or LLM-specific guidance for predictable latency.