Prompt
How do I track errors and incidents across my production app?
Latest observation
To track errors and incidents effectively in a production app, you want three layers:
- Error monitoring — catch exceptions, crashes, failed requests
- Incident alerting — notify the right people when something is broken
- Operational observability — understand what happened and how widespread it is
Here’s a practical approach.
1) Use an error tracking tool
Common options:
- Sentry
- Bugsnag
- Rollbar
- Datadog Error Tracking
- New Relic
These tools can capture:
- stack traces
- request context
- user/session info
- release version
- environment
- breadcrumbs leading up to the error
What to capture
Make sure every error event includes:
- app/service name
- environment:
prod,staging, etc. - release/build version
- timestamp
- endpoint/job name
- user or account ID if allowed
- request ID / trace ID
- severity
- tags like region, tenant, feature flag
2) Define what counts as an “incident”
Not every error is an incident. Set rules for escalation, for example:
- error rate > 5% for 5 minutes
- login failures spike
- checkout failures
- queue backlog exceeds threshold
- API latency p95 > X
- critical job failed
- a single high-value customer is impacted
This helps reduce alert noise.
3) Centralize logs, metrics, and traces
Use observability tools so you can correlate problems:
- Logs: detailed events and exceptions
- Metrics: rates, counts, latency, saturation
- Traces: request flow across services
Common stacks:
- OpenTelemetry for instrumentation
- Prometheus + Grafana
- Datadog
- Elastic Stack
- New Relic
4) Add alerting
Send alerts to:
- Slack / Microsoft Teams
- PagerDuty / Opsgenie
- email only for low-severity issues
Good alerting practices:
- alert on symptoms, not every exception
- group related errors
- deduplicate repeated events
- route by service/team/severity
- include links to logs/traces/releases
5) Build an incident workflow
When an incident happens:
- detect
- alert
- triage severity
- assign owner
- mitigate
- root-cause
- postmortem
- add a test/monitor to prevent recurrence
Use incident tracking tools like:
- PagerDuty
- Opsgenie
- Jira Service Management
- FireHydrant
- incident.io
6) Correlate errors with deployments
Track:
- deploy time
- release version
- feature flags
- config changes
This helps answer: “Did this start after the latest deploy?”
7) Add dashboards
Create dashboards for:
- error rate
- failed requests by endpoint
- latency
- crash-free sessions
- top exceptions
- incident count and MTTR
- service health
8) Use a consistent taxonomy
Example severity levels:
- SEV1: major outage or data loss
- SEV2: significant degradation
- SEV3: limited impact / workaround exists
- SEV4: minor issue / no urgent action
And categories:
- auth
- payments
- infra
- database
- third-party dependency
- deployment
- frontend
- background jobs
9) Protect user privacy
Be careful not to log:
- passwords
- access tokens
- full credit card numbers
- sensitive personal data
Redact or hash sensitive fields before sending them to monitoring tools.
10) A simple recommended setup
If you want a straightforward, effective setup:
- Sentry for error tracking
- OpenTelemetry for traces
- Datadog/Prometheus+Grafana for metrics
- PagerDuty for incident paging
- Slack for team notifications
If you want, I can also give you:
- a minimal architecture diagram
- a checklist for production readiness
- or example implementation for Node, Python, Java, or React.