Prompt
How do I add observability to an AI app in production?
Latest observation
To add observability to an AI app in production, treat it like a product + distributed system problem: you want to understand what happened, where it happened, why it happened, and what to do next.
A practical setup has 4 layers:
1) Define the key signals
For AI apps, you usually need:
- System metrics: latency, error rate, throughput, CPU/memory, queue depth, token usage, cost
- Request traces: each user request across API calls, retrieval, model calls, tools, DB access
- Model-specific data: prompt/version, model name, temperature, max tokens, output length, tool calls, citations
- Quality signals: human feedback, user retries, thumbs up/down, hallucination flags, refusal rate, task success rate
- Safety signals: toxic output, PII leakage, policy violations, jailbreak attempts
2) Instrument the app end-to-end
Add structured logging and tracing at every AI boundary:
- incoming request
- prompt construction
- retrieval step
- model invocation
- tool/function calls
- post-processing/validation
- final response
For each step, log:
request_id,user_id/tenant,session_id- model/provider
- prompt template/version
- retrieved document IDs
- latency
- tokens in/out
- cost estimate
- errors/timeouts/retries
- final outcome
Use distributed tracing so one user request becomes one trace.
3) Capture prompt and response metadata safely
You do want visibility, but be careful with sensitive data.
Best practice:
- store metadata by default
- sample or redact full prompts/responses
- hash or mask PII/secrets
- keep a secure, access-controlled audit trail for debugging
Useful fields:
- prompt template ID
- prompt hash
- model version
- system prompt version
- retrieval corpus version
- tool schema version
- moderation result
- confidence/risk score
4) Monitor quality in production
AI failures are often semantic, not technical. Add evaluation loops:
- offline evals before deployment on golden test sets
- online evals in production on sampled traffic
- human review queues for low-confidence or high-risk outputs
- feedback capture from users and operators
- drift detection on input patterns and output quality
Track:
- answer accuracy / relevance
- groundedness / citation correctness
- completion success rate
- escalation rate
- repeat question rate
- user abandonment
- model regression after prompt/model changes
A simple production observability stack
Minimum viable
- Logs: structured JSON logs
- Metrics: Prometheus / Datadog / OpenTelemetry metrics
- Tracing: OpenTelemetry traces
- Dashboards: Grafana / Datadog dashboards
- Alerting: latency spikes, error spikes, cost anomalies, safety violations
AI-specific add-ons
- prompt/version tracking
- LLM call logging wrapper
- eval service for sampled responses
- moderation/safety classifier
- feedback collection UI
- prompt diffing and release tracking
What to implement first
If you’re starting from scratch, do this in order:
- Add request IDs and distributed tracing
- Wrap every model call with structured logs
- Record token usage, latency, cost, model version
- Add prompt/version identifiers
- Log retrieval/tool call inputs and outputs
- Capture user feedback and retries
- Build dashboards and alerts
- Add offline/online evaluation pipelines
- Redact sensitive data and define retention policies
Example fields to log
A good event for one model call might look like:
{
"request_id": "req_123",
"session_id": "sess_456",
"user_id": "u_789",
"step": "llm_call",
"model": "gpt-4.1",
"prompt_template": "support_answer_v12",
"prompt_version": "12",
"retrieval_docs": ["doc_17", "doc_92"],
"latency_ms": 842,
"tokens_in": 1560,
"tokens_out": 243,
"cost_usd": 0.031,
"retry_count": 0,
"status": "success"
}
Common mistakes to avoid
- logging raw prompts/responses without redaction
- only monitoring latency and errors, not output quality
- no versioning for prompts/models/retrievers
- no linkage between user complaints and traces
- no sampling strategy, causing either data overload or blindness
- no evaluation loop for regressions after changes
Rule of thumb
If you can’t answer these questions quickly, observability is missing:
- Which model version answered this?
- What prompt and retrieved context did it use?
- How long did each step take?
- How much did it cost?
- Did it succeed from the user’s perspective?
- Was the output safe and grounded?
If you want, I can also give you:
- a reference architecture for AI observability,
- a checklist for production rollout, or
- a sample OpenTelemetry implementation in Python/Node.