Prompt
How do I monitor prompt failures, retrieval errors, and token usage together?
Latest observation
Oct 9, 2026
OpenAI APIWeb search: off
To monitor prompt failures, retrieval errors, and token usage together, the key is to treat them as one trace per user request and attach all three signal types to that trace.
1) Use a shared request/trace ID
For every incoming request:
- generate a
request_id/trace_id - pass it through:
- prompt generation
- retrieval step
- model call
- post-processing
- logging/metrics
This lets you correlate:
- prompt construction failures
- retrieval failures
- token counts / cost
- final response status
2) Emit structured events at each stage
Log each stage as a structured event with consistent fields.
Example fields:
trace_iduser_id/session_idstage=prompt | retrieval | generation | postprocessstatus=ok | errorerror_typelatency_mstokens_inputtokens_outputretrieved_docs_countretrieval_latency_msmodel_nameprompt_template_version
Example JSON:
{
"trace_id": "abc123",
"stage": "retrieval",
"status": "error",
"error_type": "vector_db_timeout",
"latency_ms": 842,
"retrieved_docs_count": 0
}
3) Track token usage on the model call
Capture:
- input tokens
- output tokens
- total tokens
- optionally estimated cost
If your provider returns usage, store it directly. If not, estimate with a tokenizer and log:
prompt_tokenscompletion_tokenstotal_tokens
4) Record failures as metrics, not just logs
Create counters/gauges like:
prompt_failure_countretrieval_error_countllm_error_countrequests_with_zero_retrieval_resultsavg_tokens_per_requestp95_total_tokenscost_per_request
Break them down by:
- model
- prompt version
- retrieval backend
- environment
- tenant/user segment
5) Build a unified dashboard
A useful dashboard usually has:
- top line: request success rate
- error breakdown: prompt vs retrieval vs model
- token usage: average, p95, per endpoint
- retrieval health: hit rate, empty result rate, timeout rate
- correlation views: errors by prompt version or retrieval backend
6) Use distributed tracing if possible
If you already use OpenTelemetry, this becomes much easier:
- one trace per request
- spans for prompt build, retrieval, LLM call
- span attributes for token usage and errors
Example span attributes:
llm.tokens.inputllm.tokens.outputretrieval.results.counterror.type
7) Practical pattern
A good implementation pattern is:
- logs for detailed event data
- metrics for aggregation and alerting
- traces for end-to-end debugging
Together, they answer:
- What failed?
- Where did it fail?
- How expensive was it?
- Which prompt/retrieval version caused it?
8) Alert on combined conditions
Examples:
- retrieval error rate > 5% for 10 minutes
- prompt failures spike after template version change
- token usage jumps 2x without traffic increase
- empty retrieval results correlate with lower answer quality
If you want, I can show you:
- a sample schema for logging these events,
- an OpenTelemetry implementation, or
- a Prometheus/Grafana dashboard design.