Prompt
How do I track token usage and latency across multiple AI apps?
Latest observation
To track token usage and latency across multiple AI apps, the cleanest approach is to treat it like observability for LLMs: instrument each app consistently, send metrics to a shared backend, and normalize them with common tags.
What to track
At minimum, capture these per request:
- Request ID
- App / service name
- Model name
- Prompt tokens
- Completion tokens
- Total tokens
- Latency
- end-to-end latency
- model/API latency
- Status
- success / error / timeout / rate-limited
- User / tenant / environment
- if applicable
Useful extras:
- Cost estimate
- Input/output size
- Cache hit/miss
- Streaming duration
- Retries
- Tool-call count
- Provider (OpenAI, Anthropic, Azure, etc.)
Recommended architecture
1. Instrument each AI app with a shared wrapper
Create a small client wrapper around your model calls that records:
- start time
- end time
- token counts from the response
- errors/retries
This wrapper should emit a standardized event like:
{
"request_id": "req_123",
"app": "support-bot",
"model": "gpt-4.1-mini",
"prompt_tokens": 1200,
"completion_tokens": 340,
"total_tokens": 1540,
"latency_ms": 842,
"status": "success",
"tenant": "acme",
"timestamp": "2026-10-10T12:34:56Z"
}
2. Send metrics to a centralized observability system
Options:
- OpenTelemetry + Grafana/Prometheus/Tempo/Loki
- Datadog
- New Relic
- Honeycomb
- OpenAI/LLM-specific tooling like LangSmith, Helicone, Arize, WhyLabs, Traceloop, Braintrust
If you want a vendor-neutral setup, OpenTelemetry is a strong default.
3. Normalize data across all apps
Use the same schema and tags everywhere:
appenvironmentmodelproviderendpointtenantuserrequest_type
This makes dashboards and alerts work across all AI apps.
How to measure latency correctly
Track more than one latency type:
- Total latency: from request start to final response
- Model latency: time spent waiting on the provider
- Queue latency: if requests are queued
- Streaming latency: time to first token and time to last token
For streaming apps, capture:
- TTFT = time to first token
- TTL = time to last token
- Tokens/sec
Token usage sources
Most AI providers return usage data directly in responses. If not:
- estimate via tokenizer libraries
- log prompt/output text lengths
- use provider-side billing/usage APIs
For accuracy, prefer provider-reported usage when available.
Best practices
- Log at the request boundary
- one record per model call
- Use correlation IDs
- trace a single user request across multiple services
- Sample intelligently
- keep full logs for errors and a subset for high-volume traffic
- Separate metrics from content
- store sensitive prompts/responses carefully
- Redact PII
- before sending to telemetry systems
- Track retries
- retries can inflate token usage and latency
- Add alerts
- token spikes, latency regressions, error-rate increases
Simple metric model
You can represent each call as:
- Counter:
llm_tokens_total{app,model,type="prompt|completion"} - Histogram:
llm_latency_ms{app,model} - Counter:
llm_requests_total{app,model,status} - Gauge/derived:
llm_cost_usd_total{app,model}
Example dashboard panels
- Tokens per app per day
- Cost per tenant
- p50/p95/p99 latency by model
- Error rate by provider
- Top endpoints by token consumption
- Time to first token for streaming apps
Practical implementation options
Option A: OpenTelemetry-based
Best if you already have infrastructure monitoring and want one standard across all apps.
Option B: LLM observability platform
Best if you want quick setup, prompt tracing, and model-specific dashboards.
Option C: Custom events to warehouse
Best if you need deep BI analysis:
- send events to Kafka/S3/BigQuery/Snowflake
- build dashboards in Looker/Metabase/Tableau
Minimal implementation pattern
- Wrap every LLM call
- Measure latency with a timer
- Extract token usage from response
- Emit one structured event
- Aggregate in a central dashboard
- Alert on anomalies
If you want, I can also give you:
- a Python example using OpenTelemetry,
- a Node.js middleware example,
- or a recommended stack for a startup vs enterprise setup.