Prompt

I'm building a production LLM workflow and need to monitor cost, latency, and errors

Technology / Observability1 observationLast seen Oct 9, 2026

Latest observation

Oct 9, 2026OpenAI APIWeb search: off

To monitor cost, latency, and errors in a production LLM workflow, set up observability at three levels: request, model, and business outcome.

1) What to measure

Cost

Track per request and aggregate:

  • Input tokens
  • Output tokens
  • Total tokens
  • Model name / version
  • Cost per request
  • Cost per user / tenant / feature
  • Cost trends over time

Useful derived metrics:

  • Cost per successful request
  • Cost per 1k requests
  • Cost per conversation/session
  • Cost by prompt template or workflow step

Latency

Track:

  • End-to-end request latency
  • Time to first token if streaming
  • Model inference latency
  • Tool / retrieval / function-call latency
  • Queue / retry delay
  • P50 / P95 / P99 latency

Break latency down by:

  • Prompt size
  • Model choice
  • Region
  • Tool usage
  • Retry count

Errors

Track:

  • API failures
  • Timeouts
  • Rate limits
  • Invalid inputs / schema mismatches
  • Tool execution failures
  • Hallucination-related downstream failures if you can detect them
  • Retry exhaustion
  • Partial stream failures

Also log:

  • Error type
  • Error message
  • HTTP status / provider status
  • Retry attempt
  • Correlation ID / trace ID

2) Instrumentation design

Add a unique request ID

Generate a correlation ID for every workflow execution and pass it through:

  • frontend
  • backend
  • LLM call
  • retrieval/tool calls
  • downstream services

This lets you trace a single user action end-to-end.

Log structured events

Use JSON logs with fields like:

  • request_id
  • user_id or tenant_id
  • workflow_name
  • model
  • prompt_tokens
  • completion_tokens
  • latency_ms
  • status
  • error_type
  • cost_usd
  • prompt_template_version

Separate spans for each step

If your workflow includes:

  • retrieval
  • reranking
  • LLM generation
  • tool execution
  • post-processing

treat each as its own measured span so you know where time and failures occur.


3) Metrics dashboard

Build a dashboard with:

  • Total cost today / this week / this month
  • Cost by model
  • Cost by feature
  • Latency percentiles
  • Error rate by provider/model
  • Token usage distribution
  • Retries per request
  • Timeouts and throttling events

Good charts:

  • Time series for cost and latency
  • Stacked bar for cost by workflow step
  • Heatmap for latency vs token count
  • Pie/bar for error categories

4) Alerting

Set alerts on:

  • Cost spikes over baseline
  • P95 latency regressions
  • Error rate > threshold
  • Rate limit spikes
  • Timeout spikes
  • Unexpected token growth
  • Model fallback frequency increasing

Example alert thresholds:

  • Cost up > 20% day-over-day
  • P95 latency > 2x trailing 7-day average
  • Error rate > 3%
  • Timeout rate > 1%
  • Retry rate > 10%

5) Tracing and observability stack

A practical stack:

  • OpenTelemetry for traces and metrics
  • Prometheus + Grafana for metrics dashboards
  • ELK / OpenSearch / Datadog / Honeycomb for logs and traces
  • A cost calculator in your app to estimate spend from token usage and provider pricing

If you use OpenAI, Anthropic, or another provider, capture:

  • model
  • usage metadata
  • finish reason
  • request duration
  • error payloads

6) Production best practices

Retry carefully

  • Retry only transient failures
  • Use exponential backoff
  • Cap max retries
  • Record retry count in metrics

Control spend

  • Set token limits
  • Choose smaller models for simple tasks
  • Cache frequent prompts/results
  • Use truncation and summarization for long contexts
  • Route to cheaper models when possible

Sample verbose logs

To reduce observability cost:

  • log everything for failures
  • sample successful high-volume traffic
  • keep full traces for a subset of requests

Redact sensitive data

Never log raw PII or secrets.

  • redact user content where necessary
  • store prompt hashes or partial content for debugging
  • encrypt logs if they may contain sensitive text

7) Minimal schema to store per request

{
  "timestamp": "2026-10-09T12:34:56Z",
  "request_id": "req_123",
  "user_id": "user_456",
  "workflow": "support_reply",
  "model": "gpt-4.1-mini",
  "prompt_tokens": 1200,
  "completion_tokens": 320,
  "total_tokens": 1520,
  "cost_usd": 0.0048,
  "latency_ms": 1832,
  "status": "success",
  "error_type": null,
  "retry_count": 0
}

8) If you want a simple implementation plan

  1. Add request_id everywhere
  2. Log every LLM call with token counts, latency, status, and model
  3. Emit metrics to Prometheus/OpenTelemetry
  4. Build Grafana/Datadog dashboards
  5. Add alerts for cost, latency, and errors
  6. Review top-cost workflows weekly
  7. Add sampling and redaction

If you want, I can also give you:

  • a reference architecture
  • a Python/TypeScript instrumentation example
  • a Grafana dashboard design
  • or a production-ready logging schema for LLM calls

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.