Prompt
How do I load test an API for latency and throughput?
Latest observation
To load test an API for latency and throughput, you want to simulate realistic traffic, measure response times under increasing load, and identify when the API starts to degrade.
1) Define what you want to measure
Common metrics:
- Latency: response time, usually reported as:
- p50 (median)
- p90 / p95 / p99
- max
- Throughput:
- requests per second (RPS)
- successful responses per second
- Also track:
- error rate
- timeouts
- CPU/memory/db usage on the server
2) Choose a load testing tool
Popular options:
- k6 — great for API testing, scripting in JS
- JMeter — GUI-based, widely used
- Locust — Python-based, easy to customize
- Gatling — strong for performance testing
- wrk / hey / ab — lightweight CLI tools for quick tests
If you want simple and modern, k6 is a strong choice.
3) Start with a baseline test
Before stressing the API, run a small test to understand normal performance:
- 1 user / low concurrency
- 1–5 minutes
- measure latency and throughput under minimal load
This gives you a comparison point.
4) Design realistic traffic patterns
Match how real clients behave:
- mix of endpoints
- realistic request sizes
- auth headers / tokens
- think time between requests
- different user flows
- read/write ratios
Avoid testing only a single “hello world” endpoint unless that’s your real use case.
5) Run different types of load tests
a) Smoke test
- Very small load
- Confirms script and environment work
b) Load test
- Gradually increase users/RPS until the system reaches expected production load
c) Stress test
- Push beyond expected load to find breaking point
d) Spike test
- Sudden jump in traffic to see how the API handles bursts
e) Soak test
- Moderate load for a long time to find memory leaks, degradation, or resource exhaustion
6) Increase load gradually
Typical approaches:
- Ramp users up over time
- Ramp RPS up over time
- Hold each step long enough to stabilize
Example:
- 10 RPS for 2 min
- 50 RPS for 5 min
- 100 RPS for 5 min
- 200 RPS for 5 min
Observe:
- latency increase
- error spikes
- saturation points
7) Measure latency correctly
Look at percentiles, not just averages.
- Average can hide bad tail latency
- p95/p99 tell you how bad it gets for slower requests
Example interpretation:
- p50 = 80 ms
- p95 = 220 ms
- p99 = 900 ms
That suggests most users are fine, but some are getting slow responses.
8) Measure throughput and saturation
Throughput usually rises with load until a bottleneck is hit, then:
- it plateaus
- latency climbs
- errors increase
This is often called the knee of the curve.
You want to know:
- maximum stable RPS
- max concurrency before degradation
- what resource becomes the bottleneck
9) Monitor the system while testing
Use application and infrastructure metrics:
- CPU
- memory
- disk I/O
- network
- DB query latency
- cache hit rate
- thread pool / connection pool usage
- queue depth
- garbage collection (if applicable)
Without server-side metrics, it’s hard to explain the results.
10) Example with k6
A simple test script:
import http from 'k6/http';
import { check, sleep } from 'k6';
export const options = {
stages: [
{ duration: '2m', target: 10 },
{ duration: '5m', target: 50 },
{ duration: '5m', target: 100 },
{ duration: '2m', target: 0 },
],
};
export default function () {
const res = http.get('https://api.example.com/items');
check(res, {
'status is 200': (r) => r.status === 200,
'latency < 500ms': (r) => r.timings.duration < 500,
});
sleep(1);
}
This ramps virtual users up over time and records response times and errors.
11) Example with wrk
Quick throughput/latency test:
wrk -t4 -c100 -d30s https://api.example.com/items
-t4= 4 threads-c100= 100 concurrent connections-d30s= 30 seconds
Good for quick benchmarking, but less flexible than k6/Locust.
12) Interpret the results
Ask:
- What load can the API handle while keeping p95 latency under the target?
- At what load does error rate rise?
- Which dependency saturates first?
- Does latency scale linearly, or does it spike after a threshold?
13) Common mistakes to avoid
- Testing only from one machine/location
- Using unrealistic data or request patterns
- Ignoring auth, DB, cache, and downstream services
- Focusing only on average latency
- Not warming up the service
- Not controlling for background traffic
- Overloading the load generator itself
14) A good basic workflow
- Pick a tool
- Define success criteria
- Create realistic test scripts
- Run a baseline
- Ramp load gradually
- Monitor app + infra
- Identify bottlenecks
- Fix and retest
If you want, I can give you:
- a k6 script for your API,
- a Locust example,
- or a step-by-step test plan for a specific endpoint.