Prompt

New Relic isn't helping me find prompt regressions, what should I use instead?

Technology / Observability1 observationLast seen Oct 9, 2026

Latest observation

Oct 9, 2026OpenAI APIWeb search: off

If New Relic isn’t giving you visibility into prompt regressions, you’ll usually want a toolchain that’s built for LLM observability + evaluation, not general APM.

Better options to use instead

1. Langfuse

Best if you want an open-source / self-hostable option.

  • Tracks prompts, model outputs, latency, cost
  • Lets you compare prompt versions
  • Supports datasets, traces, and evals
  • Good for debugging prompt changes over time

2. Arize Phoenix

Best if you want strong LLM tracing + evaluation workflows.

  • Great for prompt and response analysis
  • Has evaluation tooling for regressions
  • Good for observing hallucinations, retrieval quality, and drift
  • Strong if you’re doing RAG or agent workflows

3. LangSmith

Best if you’re already in the LangChain ecosystem.

  • Very good prompt/version tracing
  • Built-in experiment tracking and evals
  • Easy to compare runs and catch regressions
  • Strong developer UX for debugging chains/agents

4. Humanloop

Best if you want prompt management + evals + collaboration.

  • Prompt versioning
  • Regression testing
  • Human review workflows
  • Useful for teams iterating on prompts in production

5. Weights & Biases Weave

Best if your team already uses W&B or wants experiment tracking.

  • Trace LLM calls
  • Compare prompt/model versions
  • Evaluate outputs against datasets
  • Nice for experimentation-heavy workflows

If your main problem is prompt regression detection

Look for tools that support:

  • Prompt/version tracking
  • Dataset-based evaluations
  • Side-by-side comparisons
  • Automated scoring
  • Production traces linked to prompt versions
  • RAG-specific metrics if retrieval is involved

My quick recommendation

  • Open-source / self-hosted: Langfuse
  • Best for debugging and evals: Arize Phoenix
  • Best LangChain workflow: LangSmith
  • Best team prompt management: Humanloop

If you tell me:

  1. whether you’re using LangChain / LlamaIndex / custom code,
  2. whether you need self-hosted or SaaS, and
  3. whether this is chat, RAG, or agents,

…I can narrow it to the best 1–2 choices for your setup.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.