Prompt
How do I troubleshoot missing logs and traces in a SaaS app?
Latest observation
Here’s a practical way to troubleshoot missing logs and traces in a SaaS app.
1. First, narrow down the scope
Ask:
- Is it all logs/traces or only from one service / tenant / region / environment?
- Is the data missing entirely, or just delayed?
- Did the issue start after a deploy, config change, library upgrade, or infra change?
- Are logs missing but traces present, or vice versa?
That helps determine whether the issue is in:
- application instrumentation
- agent/collector pipeline
- backend ingestion
- filtering/retention/querying
2. Check instrumentation in the app
Common causes:
- Logging level too restrictive (
INFOvsDEBUG, etc.) - Trace sampling too aggressive
- Requests not instrumented due to missing middleware/interceptors
- Context propagation broken, so spans don’t link correctly
- Logging/tracing library initialization failed at startup
What to verify:
- Logs are actually emitted in the code path you expect
- Trace spans are created for that endpoint/job
- Correlation IDs / trace IDs are included if you expect them
- Any recent code changes around logging/tracing setup
3. Check the agent, sidecar, or collector
If you use an observability agent or collector:
- Is it running and healthy?
- Any restarts, crashes, OOMs, or throttling?
- Are there error messages about:
- auth failures
- network timeouts
- TLS/cert issues
- queue overflow / dropped events
- invalid payloads
Useful checks:
- Agent logs
- Collector metrics
- Health endpoints
- Pod/container status if in Kubernetes
4. Validate ingestion and transport
Missing data often happens between the app and the observability backend.
Check:
- Can the app/agent reach the SaaS endpoint?
- Any proxy, firewall, DNS, or TLS issues?
- Are events being batched and dropped under load?
- Is the payload size too large?
- Are there rate limits being hit?
If traces disappear under load, sampling or backpressure is a common culprit.
5. Review sampling settings
For traces especially, sampling is a frequent reason data “goes missing.”
Look at:
- Head-based sampling rate
- Tail-based sampling rules
- Dynamic/adaptive sampling
- Per-tenant or per-service sampling overrides
Questions:
- Are low-volume requests never sampled?
- Did someone lower the sample rate to reduce costs?
- Are errors excluded by mistake?
6. Confirm the logs are not being filtered out
Sometimes the logs exist, but your UI/query hides them.
Check:
- Time range
- Environment/service/tenant filters
- Index selection
- Severity filters
- Query syntax
- Retention window
- “Only errors” or other dashboard defaults
For traces:
- Wrong service name
- Wrong operation name
- Wrong region/account
- Filtering by tags that aren’t actually present
7. Compare with raw/local sources
If possible:
- Check local app stdout/stderr
- Check container logs
- Check host log files
- Compare with direct requests made during the same time window
This helps determine whether the issue is:
- not emitted at all
- emitted locally but not shipped
- shipped but not stored/indexed
- stored but not queryable
8. Look for schema or parsing problems
Especially in log pipelines:
- JSON malformed
- multiline logs split incorrectly
- timestamp parsing wrong
- field names changed after deploy
- trace context fields renamed or dropped
A parsing issue can make logs appear missing even though they were ingested.
9. Check quotas, limits, and retention
SaaS observability platforms may drop or delay data due to:
- ingestion quota limits
- burst limits
- retention policies
- indexing limits
- cardinality explosions
If the issue is intermittent, quota exhaustion is worth checking.
10. Use a controlled test
Create a simple test:
- emit a known log line with a unique marker
- generate a single request that should produce a trace
- verify the agent/collector sees it
- verify the backend receives it
- verify it appears in the UI
This isolates where data is being lost.
11. Common patterns and likely causes
- Logs present locally, missing in SaaS → shipper/collector/network/backend issue
- Some services missing, others fine → config inconsistency, service name mismatch, deploy drift
- Logs missing only during spikes → backpressure, buffering, rate limiting
- Traces missing but logs fine → sampling, context propagation, tracer setup
- Only one tenant missing → RBAC, tenant routing, filters, or tagging issues
12. What to collect for support/escalation
If you need vendor support, gather:
- exact time window
- environment/region
- service names
- recent deploy/config changes
- agent/collector version
- sample trace IDs / request IDs
- screenshots of filters/query
- relevant logs from app/collector
- whether it affects all tenants or one
If you want, I can also give you a step-by-step troubleshooting checklist for a specific stack like Datadog, New Relic, OpenTelemetry, Splunk, or Dynatrace.