Prompt

How do I test API rate limits and retry behavior?

Technology / Testing & QA1 observationLast seen Oct 6, 2026

Latest observation

Oct 6, 2026OpenAI APIWeb search: off

To test API rate limits and retry behavior, you want to verify two things:

  1. What happens when you send too many requests
  2. Whether your client retries correctly after being throttled or temporarily failing

Here’s a practical way to test both.


1) Understand the rate limit policy first

Check the API docs for:

  • Limit type: per second, per minute, per hour, per user, per token, etc.
  • What error code is returned: often 429 Too Many Requests
  • Headers provided:
    • Retry-After
    • X-RateLimit-Limit
    • X-RateLimit-Remaining
    • X-RateLimit-Reset

This tells you what to expect and how to assert behavior.


2) Test rate limiting by intentionally exceeding the limit

Send requests faster than the allowed threshold.

Example approach

  • If the limit is 10 requests/minute, send 15–20 requests in quick succession.
  • Observe:
    • When the API starts rejecting requests
    • Whether it returns 429
    • Whether headers indicate when to retry

Simple cURL loop

for i in {1..20}; do
  curl -i https://api.example.com/resource
done

Better: concurrent load

Use a load tool like:

  • k6
  • JMeter
  • Locust
  • hey
  • ab (ApacheBench)

This is better if the API uses short-window limits or concurrent request limits.


3) Verify retry behavior for retryable errors

Your client should usually retry only on transient failures, such as:

  • 429 Too Many Requests
  • 500 Internal Server Error
  • 502 Bad Gateway
  • 503 Service Unavailable
  • 504 Gateway Timeout

It should usually not retry on:

  • 400 Bad Request
  • 401 Unauthorized
  • 403 Forbidden
  • 404 Not Found

4) Test backoff logic

A good retry strategy usually includes:

  • Exponential backoff
  • Jitter to avoid retry storms
  • Respecting Retry-After if present
  • A max retry count

Example expected behavior

If your client gets a 429:

  • Wait for Retry-After seconds if provided
  • Otherwise wait using backoff:
    • 1s
    • 2s
    • 4s
    • 8s
  • Stop after max retries

5) Validate your client doesn’t retry too aggressively

Check that:

  • It does not retry indefinitely
  • It stops after the configured maximum
  • It doesn’t retry non-retryable errors
  • It uses increasing delays, not constant rapid retries

6) Simulate failures if you can’t hit the real limit

If the real API is hard to throttle safely, mock or stub responses.

Useful tools

  • WireMock
  • MockServer
  • Postman mock server
  • A local test proxy that returns 429/503

This lets you confirm retry logic without stressing production.


7) Check idempotency

If requests may be retried, ensure they’re safe to repeat.

  • Safe to retry:
    • GET
    • PUT if idempotent
    • DELETE if idempotent
  • Risky to retry:
    • POST unless the API supports idempotency keys

For write operations, use:

  • Idempotency-Key headers
  • Server-side deduplication

This prevents duplicate charges, orders, or records.


8) Assertions to include in tests

Your tests should verify:

  • 429 is returned when the limit is exceeded
  • Retry-After is honored
  • Retry count matches configuration
  • Backoff intervals increase correctly
  • Requests succeed after the limit resets
  • Non-retryable errors are not retried
  • Idempotent behavior prevents duplicates

9) Example test cases

Rate limit test

  1. Send requests below the limit → all succeed
  2. Send requests above the limit → some fail with 429
  3. Wait for reset window
  4. Send again → succeed

Retry test

  1. Mock 429 with Retry-After: 2
  2. Confirm client waits ~2 seconds
  3. Confirm it retries once
  4. Confirm it stops after max retries if failures continue

10) Example pseudo-logic for client retries

if response.status in [429, 500, 502, 503, 504]:
    retry up to N times
    if Retry-After header exists:
        sleep(Retry-After)
    else:
        sleep(exponential_backoff_with_jitter)
else:
    do not retry

11) Best practices

  • Test in a sandbox/staging environment
  • Avoid rate-limit tests against production unless explicitly safe
  • Use small bursts and monitor logs/metrics
  • Validate both client and server side behavior
  • Add alerts for sustained 429 rates

If you want, I can also give you:

  • a Postman test plan
  • a Python script to test rate limits
  • a pytest/unit-test example for retry logic
  • a k6 load test script

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.