Prompt
Why is my LLM latency suddenly worse in production?
Latest observation
Oct 9, 2026
OpenAI APIWeb search: off
Sudden LLM latency regressions in production usually come from one of a few buckets:
1) Model-side changes
- A different model version was deployed (even a “minor” update can change speed).
- Context window got longer: more tokens in the prompt/history means slower prefill.
- Output got longer: higher max tokens, worse stopping behavior, or more verbose prompts.
- Tool/function calling increased: extra round trips and structured-output validation can add latency.
- Safety/guardrail layers changed: moderation, policy checks, or retries can add overhead.
2) Infrastructure / serving changes
- Traffic shifted to a slower region or AZ.
- GPU/CPU resource contention: noisy neighbors, autoscaling lag, or underprovisioning.
- Cold starts / model reloads if the service scales to zero or pods recycle.
- Batching settings changed: too much batching can increase tail latency; too little can reduce throughput.
- Queueing at the load balancer or inference server due to higher QPS or smaller instance pool.
- KV cache pressure / memory fragmentation causing slower decoding.
- Network path changes: proxy, TLS, VPC peering, DNS, or egress/ingress bottlenecks.
3) Request-shape changes
- Prompt bloat: hidden system prompt growth, more conversation history, retrieved docs, or retries.
- RAG slowdown: vector DB, search, reranking, or document fetches are slower.
- More tokens per request: both input and output token counts matter.
- Concurrency increased: individual requests may be fine, but p95/p99 gets worse under load.
- Long-tail requests dominate: some users or workflows produce much larger prompts.
4) Client / application changes
- Synchronous orchestration added: waiting on multiple upstreams serially.
- Timeout/retry behavior changed, causing duplicate calls or head-of-line blocking.
- Streaming disabled or delayed: time-to-first-token may look worse if the client waits for full completion.
- Instrumentation overhead: heavy logging or tracing on the hot path.
5) Provider/vendor issues
- Model provider incident or partial degradation.
- Rate limiting / throttling leading to hidden retries or queueing.
- New regional capacity constraints.
- Policy changes or safety enforcement increasing runtime.
How to debug quickly
Compare these 5 metrics before vs after
- Input tokens
- Output tokens
- Time to first token
- Total latency
- Queue time / time waiting for model
If latency rose but token counts also rose, the model may not be slower — the requests are just bigger.
Break latency into stages
- Client prep
- Retrieval / tools
- Request queueing
- Model prefill
- Decode
- Post-processing
The slowest stage usually points to the cause.
Check deployment diffs
- Model version
- Prompt/template changes
- Max token settings
- Sampling params
- Batching/concurrency settings
- Region / instance type
- Retry logic
- RAG pipeline changes
Look at tail latencies
- p50 vs p95 vs p99
- If only p95/p99 worsened, it’s often queueing, contention, or retries.
Correlate with traffic
- Did QPS increase?
- Did request size distribution change?
- Did a new customer/workflow go live?
Fast triage checklist
- Compare old vs new token counts
- Verify model version and region
- Check queue depth, GPU utilization, OOMs, restarts
- Inspect RAG / tool call timings
- Look for retry spikes and rate limit errors
- Compare TTFT and decode speed
- Test a fixed prompt against old/new environments
Common “sudden” causes in practice
- A prompt template accidentally doubled in size
- New conversation history being appended
- Retrieval returning more documents
- Autoscaling lag after traffic spike
- Provider throttling or regional congestion
- A deployment changed max output tokens or disabled streaming
If you want, I can help you build a latency diagnosis checklist or a breakdown dashboard you can use to pinpoint the regression in a few minutes.