Prompt
What should I use to track token usage and cost per request?
Latest observation
Oct 9, 2026
OpenAI APIWeb search: off
To track token usage and cost per request, use the API response usage fields and log them alongside your request metadata.
What to capture per request
Store:
request_idor your own trace ID- model name
- input tokens
- output tokens
- total tokens
- estimated cost
- latency
- user/session ID if relevant
Where to get token usage
Most LLM APIs return a usage object in the response, for example:
prompt_tokenscompletion_tokenstotal_tokens
If you’re using a newer Responses-style API, the exact field names may vary, but the idea is the same: read the usage data from the response and persist it.
How to calculate cost
Use the model’s published pricing:
cost = (input_tokens × input_rate) + (output_tokens × output_rate)
For example:
- 1,000 input tokens at $5 / 1M tokens
- 500 output tokens at $15 / 1M tokens
Cost:
- input:
1000 / 1,000,000 × 5 = $0.005 - output:
500 / 1,000,000 × 15 = $0.0075 - total:
$0.0125
Good practice
- Centralize this in middleware/interceptors so every request is tracked
- Log raw usage first, compute cost later if pricing changes
- Separate usage by:
- model
- environment
- endpoint/use case
- tenant/customer
If you want a ready-made approach
Use:
- OpenTelemetry for traces/metrics
- a database or warehouse for usage records
- a dashboard tool like Grafana, Metabase, or Looker
Minimal schema example
{
"timestamp": "2026-10-09T12:00:00Z",
"request_id": "abc123",
"model": "gpt-4.1",
"input_tokens": 1200,
"output_tokens": 300,
"total_tokens": 1500,
"estimated_cost_usd": 0.018,
"latency_ms": 842
}
If you want, I can also give you:
- a Python example,
- a Node.js example, or
- a database schema + logging middleware pattern.