Prompt
How do I compare response times across API test runs?
Latest observation
To compare response times across API test runs, you want a process that makes the numbers consistent, comparable, and actionable.
1) Measure the same metric each time
Use the same timing definition across runs, such as:
- Average response time
- Median (p50) — often better than average if there are spikes
- p95 / p99 — useful for tail latency
- Min / max
- Throughput (requests/sec) and error rate if load differs
If you compare only averages, you may miss regressions hidden by outliers.
2) Keep test conditions as identical as possible
Response times can change for reasons unrelated to code changes. Keep these constant:
- Same endpoint and request payload
- Same environment (dev/staging/prod-like)
- Same network path/location
- Same authentication method
- Same number of virtual users / concurrency
- Same warm-up period
- Same duration and ramp-up pattern
3) Export results in a structured format
Use tools that can export run data to CSV/JSON, for example:
- Postman/Newman
- k6
- JMeter
- Locust
- Gatling
- pytest + custom timing logs
Store at least:
- timestamp
- run ID / build number
- endpoint
- response time
- status code
- payload size, if relevant
4) Compare by aggregation and distribution
For each run, calculate:
- Mean
- Median
- p90/p95
- Standard deviation
- Error rate
Then compare run A vs run B using:
- Absolute difference:
B - A - Percent change:
(B - A) / A * 100
Example:
- Run A p95 = 220 ms
- Run B p95 = 275 ms
- Change =
(275 - 220) / 220 * 100 = 25% slower
5) Use a table or chart
A simple comparison table is often enough:
| Metric | Run A | Run B | Change |
|---|---|---|---|
| Avg | 120 ms | 135 ms | +12.5% |
| Median | 110 ms | 112 ms | +1.8% |
| p95 | 220 ms | 275 ms | +25% |
| Error rate | 0.5% | 1.2% | +0.7% |
Charts help spot trends:
- Line chart over time for p95/avg
- Box plot to show distribution
- Heatmap by endpoint
6) Account for statistical noise
A difference may not be meaningful unless it’s repeated. Use:
- Multiple runs
- Confidence intervals
- A consistent sample size
If response times vary a lot, rely more on median/p95 than average.
7) Automate regression detection
Set thresholds such as:
- p95 > 10% slower than baseline
- error rate > 1%
- avg latency > 200 ms
Then fail the build or flag the run if thresholds are exceeded.
8) Best practice: compare against a baseline
Pick one run as your baseline, then compare every new run to it. That makes trend detection easier than comparing random pairs.
Quick recommendation
If you want the simplest useful approach:
- Export run results to CSV
- Compare median + p95 + error rate
- Use percent change vs a baseline run
- Plot trends over time
If you want, I can also show you:
- how to do this in Excel
- how to do it with Python/pandas
- or how to compare runs from a specific tool like Postman, JMeter, or k6.