Prompt
How do I monitor token usage, latency, and retrieval failures in one place?
Latest observation
Oct 9, 2026
OpenAI APIWeb search: off
You can monitor all three in one place by combining application-level tracing/metrics with a dashboarding/observability tool.
Best approach
Use an observability stack that captures:
- Token usage
- input/output tokens per request
- total tokens by user, endpoint, model, or session
- Latency
- end-to-end request time
- model inference time
- retrieval time
- tool/function-call time
- Retrieval failures
- no documents returned
- low similarity / low confidence
- vector DB errors
- reranker failures
- empty context passed to the model
Common setup
-
Instrument your app
- log each request with:
- request ID
- user/session ID
- model name
- prompt/completion token counts
- retrieval query
- number of docs retrieved
- retrieval score / threshold
- error status
- timestamps for each stage
- log each request with:
-
Send traces and metrics to one observability backend Good options:
- OpenTelemetry + Grafana/Tempo/Prometheus
- Datadog
- LangSmith for LLM/RAG workflows
- Arize Phoenix
- Helicone
- PromptLayer
- Lunary
-
Build one dashboard Include panels for:
- total tokens / request
- p50/p95 latency
- retrieval success rate
- retrieval error rate
- average documents retrieved
- empty-retrieval count
- top failing queries
- cost by model or route
What to track specifically for retrieval failures
A retrieval failure isn’t always a hard error, so define categories:
- Hard failure: vector DB timeout, 500s, auth failure
- Soft failure: zero docs found, low relevance scores, stale index
- Pipeline failure: retriever works, but reranker or context assembly fails
Simple metric schema
For each request, store:
request_idtimestamplatency_msinput_tokensoutput_tokensretrieval_msretrieved_countretrieval_score_maxretrieval_status(ok,empty,error,low_confidence)error_messagemodel
Practical recommendation
If you want the fastest path:
- LangSmith or Helicone for LLM/RAG request-level monitoring
- Grafana/Datadog for broader infrastructure metrics
- OpenTelemetry if you want vendor-neutral instrumentation
If you want, I can also give you:
- a minimal logging schema,
- a Grafana dashboard layout, or
- sample code for OpenTelemetry/LangSmith/Helicone.