Skip to content

Rate limits & caps

Rate limits bound how fast a key can call: requests per minute, tokens per minute, and calls in flight at once. They are separate from spend caps, which bound how much a key can spend. A rate limit answers 429 and clears by itself; a spend cap answers 402 and does not clear until its window resets or someone raises it.

Each key carries its own limits, visible in its caps on GET /v1/key. Unless your organization set other values when it created the key, they start at the defaults below.

Limit caps field Default Counted
Requests per minute rpm 60 Every request the key makes in the current clock minute, refused ones included.
Tokens per minute tpm none The worst-case tokens of each call that reaches the reservation step: an estimate of the prompt plus the full output ceiling. Counted even when the call is then refused (by this limit, a spend cap or a later check), and never corrected to the tokens actually used.
Calls in flight max_in_flight 20 Calls admitted and not yet settled — a long stream holds its slot until it ends.

Windows are clock minutes: a request at 10:00:59 and one at 10:01:00 fall in different windows. All three limits apply to sandbox keys too.

A sub-key’s limits can never be set higher than its parent’s, but they are counted per key: a sub-key’s calls do not consume its parent’s requests, tokens or in-flight slots. Spend is the opposite — see Budgets, caps & credits.

Every answer to a request whose key was admitted carries the key’s requests-per-minute window, with OpenAI’s header names:

Header Example Meaning
X-RateLimit-Limit-Requests 60 The key’s requests per minute.
X-RateLimit-Remaining-Requests 41 Requests left in the current minute.
X-RateLimit-Reset-Requests 23s Time until the window resets.

They are absent when admission itself refused the request: 401, 403 scope_denied, 400 invalid_tags or missing_tags, and the 429s of the requests window and of calls in flight. There are no token headers; read tokens_per_minute and in_flight from GET /v1/key, which needs no scope and is never refused for calls in flight.

HTTP/1.1 429 Too Many Requests
Retry-After: 23
X-Request-Id: req_6f1c0a7e2b9d4c3f8e5a1b2c3d4e5f60
{"error": {"message": "Rate limit exceeded for this key", "type": "rate_limit_error", "code": "rate_limited", "param": null}}
Code HTTP Meaning What to do
rate_limited 429 The key used up its requests per minute (“Rate limit exceeded for this key”) or its tokens per minute (“Token rate limit exceeded for this key”). Your organization may also have an organization-wide rate limit, which answers the same code. Wait Retry-After seconds — the rest of the current minute, or 60 s for the organization limit.
too_many_in_flight 429 The key already has max_in_flight calls running. Wait Retry-After (5 s), or lower your concurrency.
rate_limited 429 The model’s provider is rate limiting the model (“The provider is rate limiting this model — retry shortly”). Refused before anything was produced, the call costs nothing. Wait Retry-After (30 s), or pick another model.

The official OpenAI SDKs already retry 429 and 5xx answers — twice by default, with exponential backoff, honoring Retry-After. Raise the retry count for batch jobs; keep it low for interactive traffic. A 504 request_timeout carries x-should-retry: false, so the SDKs do not replay it: the same call would run out of time again. Send it again with stream: true instead.

Terminal window
# Retry up to 5 times on 429/503, waiting what Retry-After says.
for i in 1 2 3 4 5; do
code=$(curl -s -o out.json -D headers.txt -w '%{http_code}' https://api.belarel.com/v1/chat/completions \
-H "Authorization: Bearer $BELAREL_API_KEY" -H "Content-Type: application/json" \
-d '{"model": "<model-id>", "max_tokens": 512, "messages": [{"role": "user", "content": "Hello"}]}')
[ "$code" != 429 ] && [ "$code" != 503 ] && break
ra=$(grep -i '^retry-after:' headers.txt | tr -dc '0-9')
sleep "${ra:-5}"
done

To retry a call that might already have run — a timeout, a dropped connection — send an Idempotency-Key so the retry cannot be served twice. See Idempotency.

GET /v1/models without an Authorization header returns the public catalog. It is limited to 60 requests per minute per IP address. Its answers carry X-RateLimit-Limit and X-RateLimit-Remaining (without the -Requests suffix of the keyed routes); past the limit it answers 429 with a Retry-After header. Read the status and Retry-After rather than the body, which is not the OpenAI error shape on this one route. The public catalog may also be served from a shared cache for up to 5 minutes.

Rate limits Spend caps
Bounds Speed: requests, tokens, concurrency Money: USD per day, USD per month
Window The current minute, or calls running now The UTC calendar day and month
Refusal 429 with Retry-After 402 key_cap_exceeded
Clears By itself, within a minute When the window resets or the cap is raised
Sub-keys Counted per key Shared with the parent