Prompt
llm tracing evals cost alerts
Latest observation
Oct 9, 2026
OpenAI APIWeb search: off
Here’s a practical way to think about LLM tracing, evals, and cost alerts as one observability stack.
1) LLM tracing
Use tracing to capture what happened for each request/agent step.
What to trace
- Request/response payloads
- Prompt templates and variables
- Model name, version, temperature
- Tool calls / function calls
- Latency per step
- Token usage
- Errors / retries
- User/session IDs
- Retrieved documents / context IDs
Why it matters
- Debugging prompt issues
- Reproducing bad outputs
- Auditing agent/tool behavior
- Comparing model versions
Common tools
- OpenTelemetry
- LangSmith
- Arize Phoenix
- Helicone
- Datadog / New Relic with custom spans
2) Evals
Use evals to measure output quality automatically or with human review.
Types of evals
- Offline evals: run against a fixed dataset
- Online evals: sampled production traffic
- LLM-as-judge: another model grades outputs
- Rule-based: regex, schema validation, exact match
- Human evals: manual review for high-impact cases
What to evaluate
- Correctness
- Groundedness / citation quality
- Hallucination rate
- Tool-use accuracy
- Safety / policy compliance
- Tone / brand alignment
- JSON/schema validity
Typical workflow
- Log traces
- Sample important or failing traces
- Run evals on them
- Track scores by model/prompt/version
- Roll back or promote based on thresholds
3) Cost alerts
Use alerts to prevent token spend from growing unnoticed.
Metrics to monitor
- Total token usage
- Cost per request
- Cost per user / tenant
- Cost per workflow / endpoint
- Tool/API costs
- Average and p95 latency
- Spend by model
Alert examples
- Daily cost > $500
- Cost per request increases 20% WoW
- A tenant exceeds budget
- GPT-4 usage spikes unexpectedly
- Token usage per conversation doubles
- Latency rises while cost also rises
Best practices
- Set budgets by environment: dev/staging/prod
- Alert on both absolute cost and anomaly spikes
- Break down by model, prompt, route, and tenant
- Include estimated cost before completion if streaming
- Cap max tokens and use model fallbacks
Recommended stack
A good production setup often looks like:
- Tracing: OpenTelemetry + LLM-specific span attributes
- Storage/analysis: LangSmith, Phoenix, Datadog, or your warehouse
- Evals: automated eval pipeline + human review queue
- Alerts: Slack/PagerDuty/email from cost dashboards or metrics system
- Governance: budgets, sampling, PII redaction, access controls
A simple operating loop
- Trace every request
- Sample traces for evals
- Track quality and cost together
- Alert on regressions
- Iterate on prompts/models/tools
If you want, I can also give you:
- a reference architecture
- a tool comparison table
- or a sample implementation using OpenTelemetry + LangSmith + Slack alerts.