Rate limits & caps
Rate limits bound how fast a key can call: requests per minute, tokens per minute, and calls in flight at once. They are separate from spend caps, which bound how much a key can spend. A rate limit answers 429 and clears by itself; a spend cap answers 402 and does not clear until its window resets or someone raises it.
The limits
Section titled “The limits”Each key carries its own limits, visible in its caps on GET /v1/key. Unless your organization set other values when it created the key, they start at the defaults below.
| Limit | caps field |
Default | Counted |
|---|---|---|---|
| Requests per minute | rpm |
60 | Every request the key makes in the current clock minute, refused ones included. |
| Tokens per minute | tpm |
none | The worst-case tokens of each call that reaches the reservation step: an estimate of the prompt plus the full output ceiling. Counted even when the call is then refused (by this limit, a spend cap or a later check), and never corrected to the tokens actually used. |
| Calls in flight | max_in_flight |
20 | Calls admitted and not yet settled — a long stream holds its slot until it ends. |
Windows are clock minutes: a request at 10:00:59 and one at 10:01:00 fall in different windows. All three limits apply to sandbox keys too.
A sub-key’s limits can never be set higher than its parent’s, but they are counted per key: a sub-key’s calls do not consume its parent’s requests, tokens or in-flight slots. Spend is the opposite — see Budgets, caps & credits.
Rate limit headers
Section titled “Rate limit headers”Every answer to a request whose key was admitted carries the key’s requests-per-minute window, with OpenAI’s header names:
| Header | Example | Meaning |
|---|---|---|
X-RateLimit-Limit-Requests |
60 |
The key’s requests per minute. |
X-RateLimit-Remaining-Requests |
41 |
Requests left in the current minute. |
X-RateLimit-Reset-Requests |
23s |
Time until the window resets. |
They are absent when admission itself refused the request: 401, 403 scope_denied, 400 invalid_tags or missing_tags, and the 429s of the requests window and of calls in flight. There are no token headers; read tokens_per_minute and in_flight from GET /v1/key, which needs no scope and is never refused for calls in flight.
What a 429 looks like
Section titled “What a 429 looks like”HTTP/1.1 429 Too Many RequestsRetry-After: 23X-Request-Id: req_6f1c0a7e2b9d4c3f8e5a1b2c3d4e5f60
{"error": {"message": "Rate limit exceeded for this key", "type": "rate_limit_error", "code": "rate_limited", "param": null}}| Code | HTTP | Meaning | What to do |
|---|---|---|---|
rate_limited |
429 | The key used up its requests per minute (“Rate limit exceeded for this key”) or its tokens per minute (“Token rate limit exceeded for this key”). Your organization may also have an organization-wide rate limit, which answers the same code. | Wait Retry-After seconds — the rest of the current minute, or 60 s for the organization limit. |
too_many_in_flight |
429 | The key already has max_in_flight calls running. |
Wait Retry-After (5 s), or lower your concurrency. |
rate_limited |
429 | The model’s provider is rate limiting the model (“The provider is rate limiting this model — retry shortly”). Refused before anything was produced, the call costs nothing. | Wait Retry-After (30 s), or pick another model. |
Retry with backoff
Section titled “Retry with backoff”The official OpenAI SDKs already retry 429 and 5xx answers — twice by default, with exponential backoff, honoring Retry-After. Raise the retry count for batch jobs; keep it low for interactive traffic. A 504 request_timeout carries x-should-retry: false, so the SDKs do not replay it: the same call would run out of time again. Send it again with stream: true instead.
# Retry up to 5 times on 429/503, waiting what Retry-After says.for i in 1 2 3 4 5; do code=$(curl -s -o out.json -D headers.txt -w '%{http_code}' https://api.belarel.com/v1/chat/completions \ -H "Authorization: Bearer $BELAREL_API_KEY" -H "Content-Type: application/json" \ -d '{"model": "<model-id>", "max_tokens": 512, "messages": [{"role": "user", "content": "Hello"}]}') [ "$code" != 429 ] && [ "$code" != 503 ] && break ra=$(grep -i '^retry-after:' headers.txt | tr -dc '0-9') sleep "${ra:-5}"doneimport osfrom openai import OpenAI
client = OpenAI( base_url="https://api.belarel.com/v1", api_key=os.environ["BELAREL_API_KEY"], max_retries=5, # 429 and 5xx, honoring Retry-After)import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.belarel.com/v1", apiKey: process.env.BELAREL_API_KEY, maxRetries: 5, // 429 and 5xx, honoring Retry-After});To retry a call that might already have run — a timeout, a dropped connection — send an Idempotency-Key so the retry cannot be served twice. See Idempotency.
The public catalog
Section titled “The public catalog”GET /v1/models without an Authorization header returns the public catalog. It is limited to 60 requests per minute per IP address. Its answers carry X-RateLimit-Limit and X-RateLimit-Remaining (without the -Requests suffix of the keyed routes); past the limit it answers 429 with a Retry-After header. Read the status and Retry-After rather than the body, which is not the OpenAI error shape on this one route. The public catalog may also be served from a shared cache for up to 5 minutes.
Rate limits and spend caps
Section titled “Rate limits and spend caps”| Rate limits | Spend caps | |
|---|---|---|
| Bounds | Speed: requests, tokens, concurrency | Money: USD per day, USD per month |
| Window | The current minute, or calls running now | The UTC calendar day and month |
| Refusal | 429 with Retry-After |
402 key_cap_exceeded |
| Clears | By itself, within a minute | When the window resets or the cap is raised |
| Sub-keys | Counted per key | Shared with the parent |

