Prompt
Datadog LLM Observability feels too noisy, how do people simplify it?
Latest observation
People usually simplify Datadog LLM Observability by reducing what they trace, what they tag, and what they alert on. The goal is to keep the signal about user-facing quality and cost, not every token-level detail.
Common ways teams cut the noise
1) Trace only the important flows
Instead of tracing every LLM call:
- trace only production traffic
- sample by request type, tenant, or endpoint
- keep all traces for “important” paths like chat, agent runs, tool calls, or payment/support flows
- disable or sample down internal retries, embedding calls, or background jobs
2) Use a smaller, opinionated tag set
A lot of noise comes from high-cardinality tags.
Most teams keep only:
envservicemodelendpointorworkflowtenantonly if necessarystatus/error_type
They avoid tagging:
- raw prompts
- user IDs
- session IDs
- tool arguments
- request IDs as facets in every dashboard
- highly variable metadata
3) Stop logging full prompts/responses by default
Store full content only when needed for:
- errors
- moderation flags
- low-confidence outputs
- sampled traces
- manual review
Otherwise, keep:
- prompt length
- response length
- latency
- token counts
- cost
- success/failure
- safety/quality score
4) Focus on 4–6 metrics, not dozens
A lightweight dashboard often includes:
- request count
- p50/p95 latency
- error rate
- token usage
- cost
- quality score or “task success rate”
That usually covers most operational concerns.
5) Use spans for structure, not everything
For agentic systems, don’t instrument every tiny internal step. Instead:
- one parent span per user request
- child spans only for major steps:
- retrieval
- planning
- model call
- tool call
- post-processing
This keeps traces readable.
6) Put alerts only on user impact
Avoid alerting on every anomaly. Common alerting rules are:
- error rate spikes
- latency regression
- cost per request spike
- quality score drop
- tool failure rate increase
Teams often skip alerts on:
- minor token fluctuations
- prompt version changes
- one-off model retries
7) Use environments and release versions to segment
Instead of one giant view:
- separate dev/staging/prod
- tag by app version or prompt version
- compare only current prod vs last prod release
This makes trends easier to interpret.
8) Aggregate by workflow, not raw call
If you have many model invocations per request, aggregate them into:
- “chat answer”
- “RAG lookup”
- “summarize ticket”
- “classify intent”
That’s usually more useful than seeing 50 individual model calls.
A practical “simplified” setup
A lot of teams end up with:
Traces
- 100% of user-facing requests
- sample 10–20% of internal/background LLM calls
Tags
- service, env, model, workflow, status
Metrics
- latency, errors, tokens, cost, quality
Logs/content
- only on failures or sampled requests
Dashboards
- one overview dashboard
- one quality dashboard
- one cost dashboard
Alerts
- only on latency, error rate, quality drop, and cost spikes
Rule of thumb
If a field isn’t needed to answer one of these questions, drop it:
- Is it working?
- Is it slow?
- Is it costing too much?
- Is the output good enough?
- Which workflow/model changed?
If you want, I can sketch a minimal Datadog LLM Observability configuration for either:
- a chat app
- a RAG app
- an agentic workflow
Brands