Prompt

I need a production monitoring setup for token usage, timeouts, and errors

Technology / Observability1 observationLast seen Oct 9, 2026

Latest observation

Oct 9, 2026OpenAI APIWeb search: off

Here’s a solid production monitoring setup for token usage, timeouts, and errors for an LLM/API-based system.

1) What to monitor

Token usage

Track per request and aggregate:

  • prompt_tokens
  • completion_tokens
  • total_tokens
  • tokens/sec
  • tokens per request by endpoint / model / tenant
  • cost estimate if pricing varies by model

Timeouts / latency

Track:

  • request latency p50 / p95 / p99
  • model inference time
  • queue/wait time
  • timeout count
  • retry count
  • stream duration if using streaming responses

Errors

Track:

  • HTTP status codes
  • provider/API errors
  • rate limit errors
  • context length exceeded
  • invalid request / schema errors
  • internal exceptions
  • retry exhaustion
  • circuit breaker opens

2) Core metrics to implement

Counters

  • requests_total
  • requests_failed_total
  • timeouts_total
  • retries_total
  • token_prompt_total
  • token_completion_total
  • token_total_total
  • provider_errors_total{type=...}
  • rate_limit_errors_total
  • context_length_errors_total

Histograms

  • request_latency_seconds
  • model_latency_seconds
  • queue_wait_seconds
  • response_size_bytes
  • tokens_per_request

Gauges

  • in_flight_requests
  • current_queue_depth
  • circuit_breaker_state
  • worker_utilization

3) Logging structure

Use structured logs with these fields:

  • request_id
  • trace_id
  • user_id or tenant_id
  • model
  • endpoint
  • status
  • latency_ms
  • prompt_tokens
  • completion_tokens
  • total_tokens
  • error_type
  • error_message
  • retry_count
  • timeout_ms
  • cost_estimate

Example:

{
  "level": "error",
  "request_id": "req_123",
  "trace_id": "trace_abc",
  "tenant_id": "tenant_42",
  "model": "gpt-4.1",
  "endpoint": "/chat",
  "status": 504,
  "latency_ms": 30012,
  "timeout_ms": 30000,
  "total_tokens": 1824,
  "error_type": "timeout",
  "retry_count": 2
}

4) Tracing setup

Instrument one trace per request:

  • API ingress
  • auth/validation
  • prompt building
  • model call
  • post-processing
  • persistence / downstream calls

Add spans for:

  • prompt assembly
  • model request
  • retries
  • tool calls
  • streaming
  • fallback path

This helps separate:

  • provider latency
  • your app latency
  • network latency
  • queue delays

Use OpenTelemetry if possible.


5) Dashboards

Create a dashboard with:

Traffic

  • requests/min
  • active requests
  • queue depth

Token usage

  • total tokens over time
  • tokens/request
  • tokens by model
  • tokens by tenant
  • estimated spend

Reliability

  • error rate
  • timeout rate
  • retry rate
  • provider error breakdown

Performance

  • p50/p95/p99 latency
  • model latency
  • queue wait time

Saturation

  • worker utilization
  • concurrency limits hit
  • rate limit events

6) Alerts

Set alerts on:

  • error rate > threshold
  • timeouts spike
  • p95 latency above SLO
  • token usage spike vs baseline
  • rate limit errors increase
  • retry exhaustion
  • queue depth growing
  • cost anomaly
  • provider outage indicators

Example alert thresholds:

  • error rate > 2% for 5 min
  • timeout rate > 1% for 5 min
  • p95 latency > 10s for 10 min
  • token usage > 3x daily baseline
  • retry exhaustion > 20 in 10 min

7) Recommended stack

Simple, common stack

  • OpenTelemetry for instrumentation
  • Prometheus for metrics
  • Grafana for dashboards/alerts
  • Loki / ELK for logs
  • Jaeger / Tempo for traces

If you use a managed platform

  • Datadog
  • New Relic
  • Honeycomb
  • Grafana Cloud

8) Best practices

  • Tag metrics by model, tenant, endpoint, and environment
  • Avoid high-cardinality labels like raw user IDs in Prometheus
  • Sample logs/traces if traffic is high
  • Record both client-side and server-side timeouts
  • Distinguish soft failures (retries succeeded) from hard failures
  • Track actual billed tokens, not just estimated tokens
  • Add SLOs for:
    • availability
    • latency
    • error rate
    • cost per request

9) Minimal implementation plan

  1. Add structured logging
  2. Add OpenTelemetry traces
  3. Expose Prometheus metrics
  4. Create Grafana dashboard
  5. Add alert rules
  6. Build cost/token anomaly detection
  7. Add retry and timeout classification

10) If you want a ready-to-use schema

I can provide one of these next:

  1. Prometheus metrics definitions
  2. OpenTelemetry instrumentation example
  3. Grafana dashboard JSON
  4. Datadog monitor setup
  5. A full monitoring spec for LLM APIs

If you want, I can generate a production-ready setup for Python, Node.js, or Go.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.