Prompt
I'm building an app with serverless functions and need performance monitoring
Latest observation
Oct 8, 2026
OpenAI APIWeb search: off
For serverless apps, performance monitoring should focus on the metrics that actually affect user experience and cost:
What to monitor
- Latency
- p50 / p95 / p99 execution time
- end-to-end request latency, not just function runtime
- Cold starts
- frequency and added delay
- Errors
- function failures, timeout rates, retries
- Throughput
- invocations per minute, concurrency
- Resource usage
- memory consumption, CPU if available, duration billed
- Dependencies
- database, cache, third-party API latency and failures
- Cost signals
- invocation count, duration, storage, egress
Recommended approach
1. Instrument every function
Add structured logs and tracing:
- request ID / correlation ID
- function name
- duration
- status code / error type
- upstream dependency timings
2. Use distributed tracing
This is especially important in serverless because one user request may span: API Gateway → Function A → Queue → Function B → DB
Use:
- OpenTelemetry
- a managed APM tool if you want faster setup
3. Capture cold starts explicitly
Track:
isColdStart=true/false- initialization time
- memory size and runtime version
4. Send metrics to a time-series backend
Examples:
- CloudWatch / AWS X-Ray
- Datadog
- New Relic
- Grafana Cloud
- Honeycomb
- Prometheus + Grafana for custom setups
Good dashboards
Build dashboards for:
- function latency percentiles
- error rate by function/version
- cold start rate over time
- dependency latency
- timeout count
- invocation volume
- cost per 1k requests
Alerts to set
- error rate above threshold
- p95 latency above threshold
- timeout spikes
- cold start rate increase after deployment
- dependency failures
- sudden cost jumps
Practical tips
- Keep logs structured and consistent
- Sample traces if volume is high
- Tag metrics by environment, version, and region
- Measure after deploys to catch regressions
- Watch memory sizing: too low increases duration, too high increases cost
If you want a simple stack
A common setup is:
- OpenTelemetry for instrumentation
- Cloud provider logs/metrics for basics
- Grafana/Datadog/New Relic for dashboards and alerts
If you tell me your cloud provider and runtime, I can recommend a concrete setup and library stack.