Prompt

What should I use to track token usage and cost per request?

Technology / Observability1 observationLast seen Oct 9, 2026

Latest observation

Oct 9, 2026OpenAI APIWeb search: off

To track token usage and cost per request, use the API response usage fields and log them alongside your request metadata.

What to capture per request

Store:

  • request_id or your own trace ID
  • model name
  • input tokens
  • output tokens
  • total tokens
  • estimated cost
  • latency
  • user/session ID if relevant

Where to get token usage

Most LLM APIs return a usage object in the response, for example:

  • prompt_tokens
  • completion_tokens
  • total_tokens

If you’re using a newer Responses-style API, the exact field names may vary, but the idea is the same: read the usage data from the response and persist it.

How to calculate cost

Use the model’s published pricing:

  • cost = (input_tokens × input_rate) + (output_tokens × output_rate)

For example:

  • 1,000 input tokens at $5 / 1M tokens
  • 500 output tokens at $15 / 1M tokens

Cost:

  • input: 1000 / 1,000,000 × 5 = $0.005
  • output: 500 / 1,000,000 × 15 = $0.0075
  • total: $0.0125

Good practice

  • Centralize this in middleware/interceptors so every request is tracked
  • Log raw usage first, compute cost later if pricing changes
  • Separate usage by:
    • model
    • environment
    • endpoint/use case
    • tenant/customer

If you want a ready-made approach

Use:

  • OpenTelemetry for traces/metrics
  • a database or warehouse for usage records
  • a dashboard tool like Grafana, Metabase, or Looker

Minimal schema example

{
  "timestamp": "2026-10-09T12:00:00Z",
  "request_id": "abc123",
  "model": "gpt-4.1",
  "input_tokens": 1200,
  "output_tokens": 300,
  "total_tokens": 1500,
  "estimated_cost_usd": 0.018,
  "latency_ms": 842
}

If you want, I can also give you:

  1. a Python example,
  2. a Node.js example, or
  3. a database schema + logging middleware pattern.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.