Prompt

How do I load test an API for latency and throughput?

Technology / Testing & QA1 observationLast seen Oct 6, 2026

Latest observation

Oct 6, 2026OpenAI APIWeb search: off

To load test an API for latency and throughput, you want to simulate realistic traffic, measure response times under increasing load, and identify when the API starts to degrade.

1) Define what you want to measure

Common metrics:

  • Latency: response time, usually reported as:
    • p50 (median)
    • p90 / p95 / p99
    • max
  • Throughput:
    • requests per second (RPS)
    • successful responses per second
  • Also track:
    • error rate
    • timeouts
    • CPU/memory/db usage on the server

2) Choose a load testing tool

Popular options:

  • k6 — great for API testing, scripting in JS
  • JMeter — GUI-based, widely used
  • Locust — Python-based, easy to customize
  • Gatling — strong for performance testing
  • wrk / hey / ab — lightweight CLI tools for quick tests

If you want simple and modern, k6 is a strong choice.

3) Start with a baseline test

Before stressing the API, run a small test to understand normal performance:

  • 1 user / low concurrency
  • 1–5 minutes
  • measure latency and throughput under minimal load

This gives you a comparison point.

4) Design realistic traffic patterns

Match how real clients behave:

  • mix of endpoints
  • realistic request sizes
  • auth headers / tokens
  • think time between requests
  • different user flows
  • read/write ratios

Avoid testing only a single “hello world” endpoint unless that’s your real use case.

5) Run different types of load tests

a) Smoke test

  • Very small load
  • Confirms script and environment work

b) Load test

  • Gradually increase users/RPS until the system reaches expected production load

c) Stress test

  • Push beyond expected load to find breaking point

d) Spike test

  • Sudden jump in traffic to see how the API handles bursts

e) Soak test

  • Moderate load for a long time to find memory leaks, degradation, or resource exhaustion

6) Increase load gradually

Typical approaches:

  • Ramp users up over time
  • Ramp RPS up over time
  • Hold each step long enough to stabilize

Example:

  • 10 RPS for 2 min
  • 50 RPS for 5 min
  • 100 RPS for 5 min
  • 200 RPS for 5 min

Observe:

  • latency increase
  • error spikes
  • saturation points

7) Measure latency correctly

Look at percentiles, not just averages.

  • Average can hide bad tail latency
  • p95/p99 tell you how bad it gets for slower requests

Example interpretation:

  • p50 = 80 ms
  • p95 = 220 ms
  • p99 = 900 ms

That suggests most users are fine, but some are getting slow responses.

8) Measure throughput and saturation

Throughput usually rises with load until a bottleneck is hit, then:

  • it plateaus
  • latency climbs
  • errors increase

This is often called the knee of the curve.

You want to know:

  • maximum stable RPS
  • max concurrency before degradation
  • what resource becomes the bottleneck

9) Monitor the system while testing

Use application and infrastructure metrics:

  • CPU
  • memory
  • disk I/O
  • network
  • DB query latency
  • cache hit rate
  • thread pool / connection pool usage
  • queue depth
  • garbage collection (if applicable)

Without server-side metrics, it’s hard to explain the results.

10) Example with k6

A simple test script:

import http from 'k6/http';
import { check, sleep } from 'k6';

export const options = {
  stages: [
    { duration: '2m', target: 10 },
    { duration: '5m', target: 50 },
    { duration: '5m', target: 100 },
    { duration: '2m', target: 0 },
  ],
};

export default function () {
  const res = http.get('https://api.example.com/items');

  check(res, {
    'status is 200': (r) => r.status === 200,
    'latency < 500ms': (r) => r.timings.duration < 500,
  });

  sleep(1);
}

This ramps virtual users up over time and records response times and errors.

11) Example with wrk

Quick throughput/latency test:

wrk -t4 -c100 -d30s https://api.example.com/items
  • -t4 = 4 threads
  • -c100 = 100 concurrent connections
  • -d30s = 30 seconds

Good for quick benchmarking, but less flexible than k6/Locust.

12) Interpret the results

Ask:

  • What load can the API handle while keeping p95 latency under the target?
  • At what load does error rate rise?
  • Which dependency saturates first?
  • Does latency scale linearly, or does it spike after a threshold?

13) Common mistakes to avoid

  • Testing only from one machine/location
  • Using unrealistic data or request patterns
  • Ignoring auth, DB, cache, and downstream services
  • Focusing only on average latency
  • Not warming up the service
  • Not controlling for background traffic
  • Overloading the load generator itself

14) A good basic workflow

  1. Pick a tool
  2. Define success criteria
  3. Create realistic test scripts
  4. Run a baseline
  5. Ramp load gradually
  6. Monitor app + infra
  7. Identify bottlenecks
  8. Fix and retest

If you want, I can give you:

  • a k6 script for your API,
  • a Locust example,
  • or a step-by-step test plan for a specific endpoint.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.