Prompt
My model endpoint is timing out under load, what should I check first?
Latest observation
First, check where the timeout is happening and whether it’s caused by queueing, saturation, or the model itself.
1) Look at latency breakdown
If you have tracing/metrics, split request time into:
- client/network time
- gateway/load balancer time
- request queue time
- model inference time
- post-processing time
If queue time is growing under load, you’re likely hitting capacity limits before the model even runs.
2) Check resource saturation
On the serving host/container, inspect:
- CPU pegged?
- GPU utilization / VRAM full?
- RAM pressure / swapping?
- Disk I/O spikes?
- Open file/socket limits?
Common signs:
- High CPU but low GPU → preprocessing, tokenization, or Python overhead.
- High GPU memory / OOM retries → batch too large or model too big.
- RAM swapping → severe latency spikes and timeouts.
3) Verify concurrency and batching settings
Under load, endpoints often fail because:
- too many concurrent requests,
- batch size too large,
- no queue/backpressure,
- worker count mismatched to hardware.
Check:
- max concurrent requests
- dynamic batching window
- request queue length
- per-worker model replicas
- thread/process pool limits
4) Inspect request size and output length
Timeouts can be triggered by:
- very long prompts
- large attachments
- long generated responses (
max_tokenstoo high) - repeated retries from clients
Try comparing:
- small vs large prompt latency
- short vs long generation latency
5) Look for cold starts or reloads
If latency spikes happen intermittently:
- model is being reloaded
- autoscaling is spinning up new instances
- containers are restarting
- GPU context is being recreated
6) Check upstream/downstream dependencies
If your endpoint calls:
- feature stores
- vector DBs
- auth services
- tool/function APIs
- logging/telemetry backends
then one slow dependency can make the whole request time out.
7) Confirm timeout configuration
Make sure you know:
- client timeout
- reverse proxy timeout
- load balancer timeout
- app server timeout
- model server timeout
Sometimes the endpoint is fine but the proxy times out first.
Fastest “first check” order
- Is it queueing or inference?
- Is CPU/GPU/RAM saturated?
- Are requests too large or outputs too long?
- Are timeout settings aligned across the stack?
- Any cold starts or dependency slowness?
If you want, I can give you a 10-minute triage checklist or help you diagnose it based on your stack (e.g. vLLM, Triton, TGI, TensorRT-LLM, FastAPI, K8s, SageMaker, etc.).