Prompt
model serving latency spikes
Latest observation
Oct 9, 2026
OpenAI APIWeb search: off
Latency spikes in model serving usually come from one of a few layers: the model runtime, the hardware, the request pattern, or the serving stack around it.
Common causes
- Cold starts / autoscaling
- New pods or workers need to load model weights, warm caches, or compile graphs.
- Queueing under burst load
- Requests arrive faster than the server can batch or process them.
- Batching effects
- Dynamic batching can improve throughput but add tail latency when requests wait to fill a batch.
- GPU contention / memory pressure
- Other workloads, CUDA context switches, VRAM fragmentation, or OOM retries.
- CPU bottlenecks
- Tokenization, preprocessing, serialization, network compression, or postprocessing can dominate latency.
- Model compilation / JIT / graph warmup
- First requests may trigger tracing, kernel selection, or autotuning.
- Input variability
- Longer prompts, larger images, variable sequence lengths, or ragged tensors create tail spikes.
- GC / runtime pauses
- Python GC, JVM pauses, allocator stalls, or memory swapping.
- Networking / ingress
- Load balancer retries, TLS handshakes, cross-zone traffic, or noisy neighbors.
- Downstream dependencies
- Feature stores, vector DBs, auth services, or logging sinks can spike.
What to check first
- Break down latency by stage
- ingress → preprocessing → queue wait → inference → postprocessing → response
- Compare p50 vs p95/p99
- If p50 is stable but p99 spikes, it’s often queueing, batching, or contention.
- Correlate spikes with load and scaling events
- Deploys, scale-outs, cache misses, or node drains.
- Inspect GPU/CPU/memory metrics
- utilization, VRAM, host memory, page faults, kernel launch time.
- Look at request shape
- prompt length, batch size, image resolution, number of tokens generated.
Mitigations
- Warm pools / preloading
- Keep a few hot replicas ready.
- Limit dynamic batching delay
- Cap max queue delay to protect tail latency.
- Separate preprocessing from inference
- Move tokenization or image transforms to dedicated workers.
- Use model optimization
- Quantization, distillation, TensorRT/ONNX, KV-cache optimizations.
- Pin resources
- Dedicated GPUs, CPU affinity, isolate noisy neighbors.
- Control request sizes
- Enforce max prompt length or shard large jobs.
- Autoscaling tuned for latency
- Scale on queue depth and p95, not just CPU.
- Improve observability
- Trace each request with timestamps per stage.
If you want, I can help you diagnose this in a specific stack like Kubernetes + Triton, TorchServe, vLLM, TGI, or SageMaker.