Prompt
How do I run load tests against an API before release?
Latest observation
To run load tests against an API before release, you typically want to simulate realistic traffic and watch for latency, error rates, and resource bottlenecks under increasing load.
1) Define what “good” looks like
Before testing, set targets such as:
- Throughput: e.g. 500 requests/sec
- Latency: e.g. p95 < 200 ms, p99 < 500 ms
- Error rate: e.g. < 1%
- Resource limits: CPU, memory, DB connections, queue depth
2) Use a dedicated test environment
Don’t load test production unless you explicitly plan for it.
- Use a staging/pre-prod environment
- Match production as closely as possible:
- same API version
- similar autoscaling settings
- similar database/cache setup
- Use test data, not real user data
3) Pick a load testing tool
Common options:
- k6 – modern, scriptable, easy to CI/CD
- JMeter – mature, GUI-based, widely used
- Locust – Python-based and flexible
- Gatling – strong for JVM teams
- Artillery – simple and good for APIs
4) Model realistic user behavior
Don’t just hammer one endpoint uniformly. Include:
- authentication/login
- a mix of endpoints
- think times / pauses
- request sizes and payloads similar to real usage
- varying user roles or tenant patterns if relevant
Example traffic mix:
- 50%
GET /items - 30%
GET /items/{id} - 15%
POST /orders - 5% auth/token refresh
5) Design the test types
Usually you run several kinds:
Smoke test
Small, quick test to validate the script and endpoint behavior.
Baseline test
Low load to establish normal latency and resource usage.
Load test
Expected production traffic to verify the system can handle normal peak usage.
Stress test
Push beyond expected load to find breaking points.
Soak/endurance test
Run sustained load for hours to catch memory leaks, queue buildup, or degradation.
6) Instrument and observe the system
During the test, monitor:
- API latency and error rates
- CPU, memory, disk I/O
- network utilization
- DB performance and connection pool saturation
- cache hit rate
- thread pools / event loop saturation
- queue lag
- autoscaling behavior
Use observability tools like:
- Prometheus/Grafana
- Datadog
- New Relic
- Cloud provider metrics/logs/traces
7) Start small and ramp gradually
Example progression:
- 10 users
- 50 users
- 100 users
- 250 users
- 500 users
Hold each level long enough to stabilize, then increase.
8) Automate it in CI/CD if possible
A common approach:
- run a small smoke load test on every build
- run full load tests nightly or before release
- fail the pipeline if latency/error thresholds are exceeded
9) Analyze bottlenecks
If the test fails, look for:
- slow DB queries
- missing indexes
- inefficient serialization
- lock contention
- cache misses
- rate limiting
- too-small connection pools
- synchronous work in request paths
10) Repeat after fixes
Load testing is iterative:
- test
- identify bottleneck
- tune/fix
- test again
Example with k6
A simple k6 script:
import http from 'k6/http';
import { check, sleep } from 'k6';
import { Rate } from 'k6/metrics';
export const options = {
stages: [
{ duration: '1m', target: 50 },
{ duration: '3m', target: 200 },
{ duration: '1m', target: 0 },
],
thresholds: {
http_req_duration: ['p(95)<300'],
http_req_failed: ['rate<0.01'],
},
};
export default function () {
const res = http.get('https://staging.example.com/api/items');
check(res, {
'status is 200': (r) => r.status === 200,
});
sleep(1);
}
Run it:
k6 run test.js
Important cautions
- Don’t run tests that could hurt shared systems or real users.
- Coordinate with your ops/team if the environment is connected to real databases or third-party services.
- Make sure rate limits, alerting, and autoscaling won’t trigger unwanted side effects.
- If your API depends on external providers, mock them or isolate the dependency during testing.
If you want, I can help you with:
- a k6 script for your API,
- a Locust setup,
- or a load test plan template for release readiness.