Prompt
How do I track token usage across multiple AI apps and teams?
Latest observation
To track token usage across multiple AI apps and teams, set up a centralized usage pipeline rather than relying on each app/team reporting separately.
1) Define what you want to measure
Track at least:
- Input tokens
- Output tokens
- Model name/version
- App / service name
- Team / department
- User or tenant
- Request ID
- Cost
- Latency
- Error rate
This lets you answer:
- Which team is spending the most?
- Which app is using the most expensive model?
- Where are spikes coming from?
- What is the cost per workflow/customer?
2) Instrument every AI request
In each app, log usage at the point where the model call returns. Most model APIs return token counts in the response metadata.
Capture into a shared event schema like:
{
"timestamp": "2026-10-10T12:34:56Z",
"org": "acme",
"team": "support",
"app": "chatbot-v2",
"environment": "prod",
"model": "gpt-4.1-mini",
"request_id": "abc123",
"user_id": "u-789",
"input_tokens": 1200,
"output_tokens": 240,
"total_tokens": 1440,
"cost_usd": 0.0123,
"latency_ms": 840,
"status": "success"
}
3) Send logs to a central store
Use one of:
- Data warehouse: BigQuery, Snowflake, Redshift, Databricks
- Observability stack: Datadog, Grafana Loki, OpenTelemetry + collector
- Custom API: small internal usage service that apps post to
A warehouse is best if you want reporting, dashboards, and finance reconciliation.
4) Standardize IDs across teams
Make sure every app includes:
teamprojectcost_centerenvironmentcustomer_idortenant_idif relevant
If teams use different naming, create a mapping table:
team_alias -> canonical_teamapp_alias -> canonical_app
5) Aggregate usage by dimensions
Create scheduled queries or views that roll up:
- by day/week/month
- by team
- by app
- by model
- by environment
- by user/customer
Examples:
- tokens per team per day
- spend per model per app
- top 20 users by usage
- production vs staging spend
6) Add cost calculation
Maintain a pricing table by model and region. Then compute:
cost = input_tokens × input_rate + output_tokens × output_rate
If pricing changes, keep versioned price records so historical reports stay accurate.
7) Set budgets and alerts
Add thresholds such as:
- daily team budget
- monthly app budget
- spike detection vs moving average
- unusual usage from one user or tenant
Alert via Slack, email, PagerDuty, etc.
8) Use a proxy or gateway if possible
For better governance, route all model calls through a shared AI gateway/proxy:
- logs every request
- enforces model allowlists
- tracks cost in one place
- supports rate limits and budgets
This is often the cleanest approach in multi-team orgs.
9) Build dashboards
Useful views:
- Executive spend summary
- Team usage leaderboard
- App/model cost breakdown
- User/tenant usage
- Daily token trends
- Failed calls and retries
10) Consider privacy and access control
Token usage data can become sensitive if it includes prompts or user IDs.
- restrict prompt content logging
- mask PII
- use role-based access control
- separate finance dashboards from engineering debug logs
Recommended setup
If you want a practical default:
- Instrument every AI call
- Send usage events to a central API
- Store in a warehouse
- Join with a pricing table
- Build dashboards + alerts
- Enforce all traffic through an AI gateway over time
If you want, I can also give you:
- a database schema for token tracking,
- a Python/TypeScript logging example,
- or a reference architecture for multi-team token accounting.