Prompt
What should I use to monitor and debug LLM applications in production?
Latest observation
Oct 9, 2026
OpenAI APIWeb search: off
For monitoring and debugging LLM applications in production, you’ll usually want a combination of tools rather than just one. A good stack covers:
1. Tracing and observability
Use tools that let you inspect each request end-to-end:
- OpenTelemetry for standard tracing/metrics/logs instrumentation
- LangSmith for LLM workflow tracing, prompt/version tracking, and debugging
- Arize Phoenix for tracing, evaluation, and embedding/LLM analysis
- Helicone for LLM request logging, cost tracking, and latency monitoring
- Langfuse for observability, prompt management, traces, and evals
2. Error and performance monitoring
For app-level issues:
- Sentry for exceptions, stack traces, and release monitoring
- Datadog / New Relic / Grafana Cloud for infrastructure and application metrics
3. Evaluation and quality monitoring
To catch regressions in outputs:
- Automated eval frameworks like Ragas, DeepEval, or TruLens
- Custom evaluation pipelines for groundedness, relevance, toxicity, hallucination rate, etc.
4. Logging
Keep structured logs of:
- prompts, model responses, tool calls
- token usage, latency, model name/version
- user/session IDs and request metadata
Be careful to redact sensitive data.
What I’d recommend by use case
If you’re building with LangChain/LlamaIndex and want fastest debugging
- LangSmith or Langfuse
If you want strong open-source observability
- OpenTelemetry + Langfuse + Phoenix
If you care mostly about cost/latency/request analytics
- Helicone or Langfuse
If you want enterprise APM plus LLM support
- Datadog or New Relic combined with OpenTelemetry
Practical minimum setup
If you want a solid production baseline, use:
- OpenTelemetry for traces/metrics
- An LLM observability platform like Langfuse or LangSmith
- Sentry for exceptions
- Automated evals in CI and production sampling
Key things to monitor
- latency per request and per tool call
- token usage and cost
- error rates and retries
- retrieval quality
- prompt/version changes
- hallucination/grounding issues
- user feedback and conversion metrics
If you want, I can also suggest a specific stack for Python, Node.js, or AWS/GCP, or compare Langfuse vs LangSmith vs Helicone.