Prompt
LLM app observability tools
Latest observation
Oct 9, 2026
OpenAI APIWeb search: off
Here are some of the main LLM app observability tools people use today, grouped by what they’re good at:
Open-source / self-hosted
- Langfuse — Tracing, prompt management, evaluations, user feedback, cost/latency tracking.
- OpenLIT — OpenTelemetry-native observability for LLM apps; traces, metrics, logs.
- Phoenix (Arize Phoenix) — Tracing, evaluations, RAG analysis, debugging.
- Helicone — LLM gateway + observability; request logging, caching, rate limiting, analytics.
- Traceloop — OpenTelemetry-based tracing for LLMs and agents.
- AgentOps — Agent-focused observability and debugging.
Commercial platforms
- LangSmith — From LangChain; tracing, evals, datasets, prompt testing, debugging.
- Arize AI — Model/LLM monitoring, evals, drift, production monitoring.
- Datadog LLM Observability — Integrates with existing infra monitoring; traces, logs, metrics.
- New Relic — APM + LLM observability integration.
- Honeycomb — High-cardinality tracing/observability for LLM workflows.
- WhyLabs — Monitoring, drift, data quality, LLM/app observability.
- Braintrust — Evals, tracing, prompt experiments, feedback loops.
- Humanloop — Prompt/versioning, evals, human feedback, monitoring.
- PromptLayer — Prompt tracking, versioning, analytics.
- Galileo — Eval-focused platform for LLM and RAG quality.
Useful comparison dimensions
When choosing, compare:
- Tracing depth: do you need step-by-step agent/tool traces?
- Eval support: offline test sets, online feedback, automated judge evals
- RAG visibility: retrieval quality, chunk relevance, grounding
- Prompt management: versions, rollbacks, experiments
- Cost tracking: token usage, model/provider breakdown
- OpenTelemetry support: important if you already use standard observability tooling
- Deployment model: SaaS vs self-hosted
- PII/security: redaction, data retention, compliance
Quick recommendations
- Building with LangChain → LangSmith or Langfuse
- Want open-source + flexible → Langfuse, Phoenix, OpenLIT
- Need strong evals → LangSmith, Braintrust, Galileo
- Already on Datadog/New Relic → use their LLM observability modules
- Want an LLM gateway → Helicone
If you want, I can also give you:
- a feature comparison table,
- recommendations by stack/use case (RAG, agents, chatbots, internal tools), or
- a minimal observability setup for your app.