Prompt
How do I track token usage and spend across multiple AI apps?
Latest observation
To track token usage and spend across multiple AI apps, set up a central usage accounting layer instead of relying on each app’s dashboard separately.
1) Standardize logging in every app
Have each app log the same fields for every model call:
app_nameenvironment(prod/dev)user_id/team_id/project_idprovider(OpenAI, Anthropic, etc.)modelinput_tokensoutput_tokenscached_tokens/reasoning_tokensif availablerequest_idtimestampcost_usd(calculated or stored)latency_msstatus/ error info
If you can’t get token counts directly from the SDK response, estimate them with the provider’s tokenizer or compute them after the call using prompt + completion text.
2) Create a shared cost calculator
Maintain a pricing table by provider/model and calculate cost centrally:
- Prompt tokens × input rate
- Output tokens × output rate
- Add special handling for:
- cached input tokens
- audio/image tokens
- tool calls / function calling
- fine-tuning / embedding models
- batch discounts if applicable
Keep the price table versioned, because model pricing changes.
3) Send all events to one place
Use one of:
- Database: Postgres + dashboards
- Analytics pipeline: BigQuery/Snowflake + BI tool
- Observability stack: OpenTelemetry, Datadog, Grafana, etc.
- Dedicated AI cost tooling: internal or third-party usage dashboards
A simple architecture:
- App SDK wrapper logs each call
- Events go to Kafka/queue/webhook
- Worker enriches with pricing
- Store in warehouse/DB
- Dashboard shows usage by app/team/user/model/day
4) Enforce per-app or per-team budgets
Add budget controls:
- daily/monthly spend limits
- warnings at 50/80/100%
- per-user quotas
- environment-specific caps
- rate limits on expensive models
This helps you see spend and prevent overruns.
5) Use request correlation
Assign a unique ID to each generation request and propagate it through:
- frontend
- backend
- retries
- tool calls
- sub-requests
This avoids double counting and helps trace spend per feature.
6) Track at multiple dimensions
Useful breakdowns:
- by app
- by team
- by customer
- by feature
- by model
- by environment
- by time window
- by request type (chat, summarization, embedding, rerank)
7) Watch for common gotchas
- Retries can inflate usage if you log every attempt
- Streaming responses still need final token totals
- Caching may reduce cost but needs explicit tracking
- Multiple providers have different billing units
- Tool/function calls may not be obvious from token counts alone
- Some providers bill for “reasoning” tokens separately
8) If you want a simple MVP
Start with:
- A shared wrapper around all model calls
- Log token counts and model name
- Join with a pricing table
- Store in Postgres
- Build a basic dashboard in Metabase/Grafana
9) Example schema
ai_usage_events (
id,
request_id,
app_name,
team_id,
user_id,
provider,
model,
input_tokens,
output_tokens,
cost_usd,
timestamp,
status
)
10) Example calculation
If a call uses:
- 1,200 input tokens at $3 / 1M
- 400 output tokens at $15 / 1M
Cost:
- Input: 1,200 × 3 / 1,000,000 = $0.0036
- Output: 400 × 15 / 1,000,000 = $0.0060
- Total: $0.0096
If you want, I can also help you design:
- a database schema
- a token/cost logging middleware
- a dashboard layout
- or a provider-agnostic cost calculator in Python/TypeScript.