Prompt
I'm building a production LLM app and need observability, cost tracking, and rollback
Latest observation
For a production LLM app, you want three things working together:
- Observability — know what the model is doing
- Cost tracking — know what each request/user/workflow costs
- Rollback — safely revert prompts/models/tools/configs when quality drops
Here’s a practical setup.
1) Observability: what to capture
Track every LLM request as a structured event.
Core fields
- request_id / trace_id
- user_id / tenant_id
- timestamp
- app version / prompt version / model version
- provider (OpenAI, Anthropic, local, etc.)
- model name
- input tokens / output tokens / total tokens
- latency
- status (success/error/timeout)
- error type
- tool calls and tool latency
- retrieval stats if using RAG:
- documents retrieved
- top-k scores
- sources used
- safety flags / moderation results
- quality signals:
- user rating
- thumbs up/down
- abandonment
- task success/failure
Log at multiple levels
- Request logs: one row per call
- Span tracing: prompt build → retrieval → model call → tool call → post-processing
- Conversation logs: message history for debugging
- Evaluation logs: offline benchmark runs and regression tests
What to avoid
- Logging raw sensitive data by default
- Logging secrets, API keys, PII, payment data
- Keeping indefinite full prompt histories without access controls
2) Cost tracking: attribute spend correctly
You need cost to be attributable by:
- user
- org/tenant
- endpoint
- feature
- model
- environment (dev/staging/prod)
- experiment/canary
Cost calculation
Usually:
input_tokens * input_rate + output_tokens * output_rate- plus tool costs
- plus embedding/retrieval/vector DB costs
- plus image/audio costs if applicable
Recommended approach
Maintain a cost table per model/provider, then compute per request.
Example schema:
modelinput_cost_per_1k_tokensoutput_cost_per_1k_tokenseffective_date
Then emit:
request_cost_usdestimated_cost_usdif provider billing is delayedbilled_cost_usdwhen reconciled
Cost controls
- Per-tenant budgets
- Rate limits
- Max tokens per request
- Automatic downgrade to cheaper model when confidence is high enough
- Caching for repeated prompts/results
- Batch non-urgent jobs
Useful metrics
- cost per successful task
- cost per conversation
- cost per active user
- cost per feature
- cost by model and prompt version
3) Rollback: make every change reversible
In LLM apps, “rollback” should apply to:
- prompt templates
- model choice
- tool schemas
- system instructions
- retrieval settings
- safety thresholds
- post-processing logic
Best practice: version everything
Treat these as deployable artifacts:
prompt_v12tool_schema_v5rag_config_v3model_config_v8
Use:
- Git for source control
- immutable artifact storage for releases
- config flags or feature flags for runtime selection
Rollback mechanisms
A. Feature flags
Route traffic by:
- percentage
- tenant
- user segment
- geography
- environment
If quality drops, flip traffic back to the previous version.
B. Canary release
Send:
- 1–5% traffic to new prompt/model
- compare metrics against baseline
- promote only if metrics improve
C. Shadow mode
Run the new version in parallel without showing results to users. Compare outputs and metrics before cutover.
D. Version pinning
Every request should record:
- prompt version
- model version
- tool version
- retrieval config version
That way you can reproduce bad behavior and rollback precisely.
4) Recommended architecture
A solid production setup looks like this:
Request flow
- API receives user request
- Create
trace_id - Load config by feature flag/version
- Build prompt
- Retrieve context if needed
- Call model
- Execute tools if needed
- Post-process output
- Log:
- tokens
- latency
- cost
- errors
- quality signals
- Send metrics to dashboards/alerts
Data pipeline
- Application logs → central logging
- Tracing → OpenTelemetry-compatible backend
- Metrics → Prometheus/Grafana or Datadog
- Analytics → warehouse like BigQuery/Snowflake/Postgres
- Experiment tracking → A/B testing or ML eval platform
5) Alerts you should set up
Alert on:
- error rate spike
- latency spike
- token usage spike
- cost spike
- output length anomaly
- tool failure spike
- retrieval empty-rate increase
- safety violation increase
- user thumbs-down rate increase
- conversation abandonment spike
6) Practical stack options
Lightweight stack
- OpenTelemetry for tracing
- Postgres for request/cost logs
- Grafana/Prometheus for metrics
- Sentry for errors
- Feature flags: LaunchDarkly, Unleash, or homegrown
More managed stack
- Datadog / New Relic / Honeycomb for traces + metrics
- BigQuery/Snowflake for analytics
- LaunchDarkly for rollback and experimentation
- OpenAI/Anthropic usage APIs or internal metering for billing reconciliation
LLM-specific tools
Depending on your stack:
- LangSmith
- Helicone
- LiteLLM
- Arize Phoenix
- W&B Weave These can help with tracing, evals, and debugging model behavior.
7) A simple data model
At minimum, store a llm_request record like:
{
"trace_id": "abc123",
"user_id": "u_42",
"tenant_id": "t_9",
"feature": "support_agent",
"prompt_version": "prompt_v12",
"model": "gpt-4.1-mini",
"provider": "openai",
"input_tokens": 1200,
"output_tokens": 340,
"latency_ms": 1820,
"request_cost_usd": 0.0142,
"status": "success",
"error": null,
"retrieved_docs": 4,
"tool_calls": 1,
"thumbs_up": null,
"created_at": "2026-10-09T12:34:56Z"
}
8) Rollback playbook
When something regresses:
- Detect via alert or eval
- Check which version changed
- Compare to baseline:
- cost
- latency
- success rate
- user ratings
- Disable canary / flip feature flag
- Revert to prior prompt/model/tool config
- Preserve bad traces for debugging
- Run postmortem and add test coverage
9) Minimum viable production checklist
You should have:
- Request tracing
- Prompt/version logging
- Token + cost metering
- Per-tenant budgeting
- Dashboards for latency/cost/error rate
- A/B or canary release system
- One-click rollback via flags
- Offline evals before promotion
- PII/sensitive data redaction
- Audit logs and access control
If you want, I can also give you:
- a reference architecture diagram,
- a Postgres schema for observability/cost tracking, or
- a production rollout checklist for prompts/models.