Prompt

How do I set up alerts for latency, token spend, and failure spikes in an LLM app?

Technology / Observability1 observationLast seen Oct 9, 2026

Latest observation

Oct 9, 2026OpenAI APIWeb search: off

Here’s a practical way to set up alerts for an LLM app so you catch problems early without getting paged for noise.

1) Track the right metrics

At minimum, emit these per request:

Latency

  • request_latency_ms
  • Break out by:
    • model
    • endpoint / route
    • customer / tenant
    • success vs error
    • streaming vs non-streaming

Useful aggregations:

  • p50, p95, p99 latency
  • average latency
  • tail latency by model/version

Token spend

  • prompt_tokens
  • completion_tokens
  • total_tokens
  • estimated_cost_usd

Also track:

  • tokens per request
  • tokens per user / org / day
  • spend by model, route, and tenant

Failures

  • request_count
  • error_count
  • timeout_count
  • rate_limit_count
  • invalid_response_count
  • tool_call_failure_count

Calculate:

  • error rate = errors / total_requests
  • timeout rate
  • 429 rate
  • retry rate

2) Use alert thresholds that match the metric

A good alerting setup usually has:

  • absolute thresholds for hard limits
  • relative anomaly alerts for spikes
  • burn-rate alerts for sustained issues

Latency alerts

Examples:

  • p95 latency > 2s for 10 minutes
  • p99 latency > 5s for 5 minutes
  • latency increased 2x vs 7-day baseline

A good pattern:

  • Warning: p95 above threshold for 10–15 min
  • Critical: p99 above threshold or big regression vs baseline

Token spend alerts

Examples:

  • Daily spend > 80% of budget
  • Hourly spend > 2x normal rate
  • Tokens per request increased 50% vs baseline
  • Completion tokens spike on a specific route/model

If you have a budget:

  • Alert at 50%, 80%, 90%, and 100%
  • Add anomaly alerts for sudden growth

Failure spike alerts

Examples:

  • Error rate > 5% for 5 minutes
  • Timeout rate > 2% for 10 minutes
  • 429 rate > 1% for 5 minutes
  • Success rate drops below 95%

For production, a classic SRE-style alert is:

  • 5xx/total > threshold over rolling window
  • burn-rate alert for your error budget

3) Alert on ratios, not just counts

Counts alone are noisy. Prefer:

  • error_rate = errors / requests
  • timeout_rate = timeouts / requests
  • token_cost_per_request
  • latency_p95 compared to baseline

This avoids false alarms when traffic is low or high.

Example:

  • Don’t alert on “100 failures”
  • Do alert on “failure rate above 3% for 10 minutes, with at least 200 requests”

That minimum volume guard is important.


4) Add baseline/anomaly detection for token spend and latency

LLM apps often have natural variability, so static thresholds aren’t enough.

Use:

  • moving average
  • day-over-day comparison
  • week-over-week comparison
  • seasonal baselines by hour/day

Good anomaly examples:

  • token spend today is 1.8x the same hour last week
  • p95 latency is 2 standard deviations above normal
  • completion length jumped after a prompt/model change

5) Segment alerts by model, route, and tenant

Most LLM incidents are localized.

Break alerts down by:

  • model version
  • prompt template
  • feature flag
  • customer / tenant
  • geography / region
  • provider (OpenAI, Anthropic, local model, etc.)

This helps you avoid broad alerts when only one route is broken.


6) Create a small alert matrix

A simple starting point:

Latency

  • Warning: p95 > 1500 ms for 10 min
  • Critical: p99 > 4000 ms for 5 min

Spend

  • Warning: daily spend > 75% of budget
  • Critical: daily spend > 90% of budget
  • Spike: spend/hour > 2x 7-day hourly baseline

Failures

  • Warning: error rate > 2% for 10 min
  • Critical: error rate > 5% for 5 min
  • Spike: timeout rate > 1% for 5 min
  • Spike: 429 rate > 1% for 5 min

Adjust the numbers based on your traffic and business tolerance.


7) Instrument before you alert

Alerts are only useful if the underlying data is reliable.

Log at least:

  • request ID
  • user/tenant ID
  • model name/version
  • prompt template version
  • token counts
  • latency breakdown:
    • queue time
    • provider latency
    • tool latency
    • post-processing latency
  • response status / error type

If possible, also capture:

  • retries
  • cached vs uncached requests
  • streaming duration
  • prompt length and output length

8) Use dashboards plus alerts

Before creating alerts, build dashboards for:

  • total requests
  • latency percentiles
  • token spend over time
  • errors by type
  • spend by model/tenant
  • top slow routes
  • top expensive prompts

Then create alerts from the same metrics.


9) Route alerts by severity

Example:

  • Warning → Slack / email
  • Critical → PagerDuty / phone
  • Business spend breach → finance/ops channel too

Also include:

  • current value
  • baseline
  • impacted route/model
  • link to dashboard
  • recent deploy or prompt change if available

10) Reduce noisy alerts

Best practices:

  • require N minutes over threshold
  • require minimum traffic
  • deduplicate similar alerts
  • suppress during planned maintenance
  • use different thresholds for dev/staging/prod
  • alert only on user-facing paths, not every background job

Example alert rules

Latency

  • p95(request_latency_ms) > 2000 for 10m
  • p99(request_latency_ms) > 5000 for 5m

Token spend

  • sum(estimated_cost_usd) over 1h > 2x avg of same hour last 7 days
  • daily_spend > budget * 0.8

Failure

  • error_rate > 0.05 for 5m and request_count > 200
  • timeout_rate > 0.02 for 10m

Suggested stack

Common implementation:

  • Metrics: Prometheus, Datadog, OpenTelemetry
  • Dashboards: Grafana, Datadog dashboards
  • Alerts: PagerDuty, Opsgenie, Slack
  • Logs/traces: OpenTelemetry + your logging backend

If you want, I can also give you:

  1. a Prometheus alert rule set for these metrics,
  2. a Datadog monitor setup, or
  3. a Python/Node instrumentation example for an LLM app.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.