Errors & debugging
The API speaks the OpenAI error dialect, so an official OpenAI SDK reads Belarel errors natively. This page lists the codes you can receive, what each means, and what to do about it.
The error envelope
Section titled “The error envelope”Every error from a /v1 endpoint has this shape:
{ "error": { "message": "Model \"<model-id>\" is not among the models this organization opened to its API keys", "type": "permission_error", "code": "model_not_allowed", "param": "model" }, "belarel": { "request_id": "req_6f1c0e2a9b8d4c71a3e5f0b2d4c6e8a0", "policy": { "decisions": [{ "rule": "org_api_restriction", "verdict": "deny" }] } }}| Field | Meaning |
|---|---|
error.code |
Stable, machine-readable code. Branch on this, never on message. |
error.type |
The OpenAI error class, derived from the HTTP status (see below). |
error.message |
A human-readable explanation. Wording can change. |
error.param |
The request field at fault, when there is one — model, max_tokens, stop, Idempotency-Key… — else null. |
belarel |
Present on policy refusals, on 402 max_cost_per_request and on 404 model_not_found: the request id and the policy rule that decided (unknown_model on a 404). SDKs ignore this extra key. |
error.type follows the status: 400, 405, 422 → invalid_request_error · 401 → authentication_error · 402 → insufficient_quota · 403 → permission_error · 404 → not_found_error · 409 → conflict_error · 429 → rate_limit_error · 5xx → api_error.
Two rules to build on
Section titled “Two rules to build on”- A refusal is a
4xx. Money, policy, parameter and rate refusals all come back as4xx, and they name what refused the call. - A
5xxmeans the API could not decide, could not reach the model, or ran out of time — so it refused. It never serves a call it could not check. Most5xxanswers are worth retrying afterRetry-After. Four are not, because they answer the same until something changes:501 not_implemented,503 pricing_unavailable, a503 admission_unavailablecaused by an incomplete BYOK setup, and504 request_timeout(retry it withstream: trueinstead).
Policy refusals
Section titled “Policy refusals”When your organization’s policy refuses a model, the response names the rule in belarel.policy.decisions[], with verdict: "deny":
| Rule | Why the model was refused |
|---|---|
contract_blocked_models |
Your contract blocks this model. |
contract_allowed_models |
Your contract allows a list of models, and this one is not on it. |
contract_allowed_providers |
The model is served by a provider your contract does not allow. |
residency |
The model is not available in your organization’s residency zone. |
org_api_restriction |
Your organization limited the models open to API keys, and this one is not among them. |
org_api_disabled |
Your organization turned off this kind of model for API keys. |
key_allowed_models |
This key is restricted to a list of models, and this one is not on it. |
contract_max_cost_per_request |
The worst-case estimated cost of the call exceeds your contract’s maximum per request. |
A model your organization cannot see at all is never a 403: it answers 404 model_not_found, exactly like a model that does not exist, with the generic decision {"rule": "unknown_model", "verdict": "deny"}.
Not every 403 model_not_allowed carries decisions: a refusal decided at the reservation step, or by a last availability check just before the model is called, comes without a belarel block.
Error codes
Section titled “Error codes”Retry: Yes — retry after Retry-After (or with backoff). No — retrying the same request gives the same answer; change the request or the configuration first.
400 — the request cannot be served as sent
Section titled “400 — the request cannot be served as sent”| Code | Meaning | Retry? | What to do |
|---|---|---|---|
invalid_request |
A field is missing or invalid — a wrong type, temperature outside 0–2, top_p outside 0–1, a token budget that is not a positive integer, more than 4 stop sequences, a seed that is not an integer. Also used by /v1/embeddings, /v1/usage and /v1/keys bodies. param names the field. |
No | Fix the field named in param. |
invalid_request |
The model’s provider refused the request as written. The message starts with “The provider refused the request:” followed by the provider’s explanation, cleaned of links and secrets. | No | Fix the request; do not resend it unchanged. |
unsupported_parameter |
A parameter this API does not support, named in param. See Unsupported parameters. |
No | Remove the parameter. |
validation_failed |
Chat Completions without model. |
No | Send model. |
invalid_messages |
messages is not a non-empty array of {role, content}, or a message uses a role other than system, user, assistant or tool — the developer role included. |
No | Fix messages; send developer instructions as system. |
invalid_input |
Responses input could not be read as a conversation. |
No | Fix input. |
invalid_tags |
A tag is malformed: bad key, value over 200 characters, more than 16 tags, or metadata is not an object of strings. |
No | Fix the tags. |
missing_tags |
A model call (chat, responses, embeddings) did not send a tag the key requires. The message lists them. Reads never need tags. | No | Add the required tags. |
max_tokens_exceeds_key_limit |
The output budget is above this key’s limit. param is max_tokens on Chat Completions — even when you sent max_completion_tokens — and max_output_tokens on Responses. |
No | Lower it, or omit it to use the key’s limit. |
model_lacks_vision |
The request contains images and the model cannot read them. | No | Pick a vision-capable model. |
model_lacks_tools |
The request declares tools, and the model cannot call tools. |
No | Pick a model with the tools capability, or drop tools. |
model_lacks_structured_output |
response_format / text.format asks for structured output the model cannot produce. param is always response_format, on Responses too. |
No | Pick a model with the structured_output capability. |
web_search_unavailable |
Web search was requested on a model with no native web search. | No | Pick another model, or drop web search. |
connectors_require_user_session |
A non-empty connectors or hitl_approvals was sent; they are not available to API keys. |
No | Remove them. |
model_not_zdr |
A zero_retention key called a model that cannot be served with zero retention. |
No | Pick a model with zero_retention_available: true. |
byok_unavailable |
A BYOK key called a model (or embeddings) that cannot run on your own provider key. | No | Pick a model with byok_available: true. |
byok_credential_rejected |
A BYOK key, and the provider refused your organization’s own provider key. | No | Check or replace the provider key in your organization’s BYOK settings. |
model_not_allowed |
On /v1/embeddings: a model other than the one embedding model served. |
No | Use the model listed for embeddings. |
child_exceeds_parent |
A sub-key would exceed its parent (scopes, models, caps, rates, expiry). | No | Tighten the sub-key’s settings. |
model_not_in_catalog |
A sub-key names models outside your organization’s API catalog. | No | Use models from GET /v1/models. |
key_revoked |
You tried to disable or re-enable a revoked sub-key. | No | Create a new sub-key. |
Unsupported parameters
Section titled “Unsupported parameters”On Chat Completions and Responses, every top-level request field is either forwarded to the model, refused by name, or listed below as having no effect. A top-level field the endpoint does not know is refused with unsupported_parameter, like OpenAI does. So are these known fields:
- Chat Completions:
nabove 1,logprobs: true,top_logprobs, a non-emptylogit_bias,store: true,prediction,audio,modalitiesother than text, atool_choicethat forces a tool (onlyautoandnoneare accepted), and the legacyfunctions/function_call. - Responses:
store: true,previous_response_id,item_referenceandinput_fileinputs, tool types other thanfunctionandweb_search, atool_choiceother thanautoornone,truncation: "auto",background: true,top_logprobs.
Some fields are accepted and have no effect: user, stream_options, parallel_tool_calls and service_tier on Chat Completions; user, parallel_tool_calls and service_tier on Responses, where stream_options is refused with unsupported_parameter.
The check stops at the top level. Fields nested inside an accepted object — for example text.verbosity or reasoning.summary on Responses, or json_schema.strict inside response_format — are not checked: one the API does not use is ignored without an error. /v1/embeddings has no such check at all: it reads model, input, encoding_format and dimensions, and ignores any other field.
A forwarded parameter that the model’s provider then ignores (some ignore seed, the penalties, or top_p next to temperature) is not an error: the call succeeds and lists it in belarel.ignored_parameters. That list uses the Chat Completions names on both endpoints: on Responses, an ignored max_output_tokens is listed as max_tokens, and an ignored text.format as response_format.
401 — authentication
Section titled “401 — authentication”| Code | Meaning | Retry? | What to do |
|---|---|---|---|
invalid_api_key |
No key, a bearer that is not a Belarel key, or an unknown, revoked, expired or disabled key — deliberately one answer, on every endpoint. The one exception: GET /v1/models with no Authorization header answers 200 with the public catalog. |
No | Send Authorization: Bearer bel_…, and check the key’s status in your dashboard. |
402 — money
Section titled “402 — money”| Code | Meaning | Retry? | What to do |
|---|---|---|---|
key_cap_exceeded |
The key’s daily or monthly spend cap — or its parent’s — is reached. | No | Wait for the window to reset, or raise the cap. A webhook can alert you. |
insufficient_credits |
Your organization does not have enough credits for this call. | No | Top up, or check your subscription. |
budget_exceeded |
Your organization’s budget for this period is exhausted. | No | Ask an administrator to review the budget. |
spend_refused |
Your organization’s spend controls refused the call. | No | Ask an administrator. |
max_cost_per_request |
The worst-case estimated cost is above your contract’s per-request maximum. | No | Lower max_tokens. |
403 — permission
Section titled “403 — permission”| Code | Meaning | Retry? | What to do |
|---|---|---|---|
scope_denied |
The key lacks the scope this endpoint needs. | No | Use a key with the scope. |
model_not_allowed |
The model is not allowed. On a policy refusal, belarel.policy.decisions names the rule. |
No | Pick another model, or change the policy. |
model_not_priced |
The model has no price yet and cannot be served. | No | Pick another model. |
web_search_not_allowed |
Web search is not available for this model in your organization. | No | Drop web search, or pick another model. |
entitlement_denied |
Your organization’s plan does not allow this call. | No | Contact your administrator. |
model_ambiguous |
The model id could not be resolved to a single model. Rare. | No | Contact support with the X-Request-Id. |
404, 405 — not found
Section titled “404, 405 — not found”| Code | HTTP | Meaning | Retry? | What to do |
|---|---|---|---|---|
model_not_found |
404 | The model does not exist or is not available to your organization. Carries the decision unknown_model. |
No | List models with GET /v1/models. |
not_found |
404 | No such request (/v1/requests/{id}) or sub-key (/v1/keys/{id}) for this key. |
No | Check the id and the key you use. |
method_not_allowed |
405 | Wrong HTTP method for this endpoint. | No | Use the documented method. |
409, 422 — conflicts
Section titled “409, 422 — conflicts”| Code | HTTP | Meaning | Retry? | What to do |
|---|---|---|---|---|
idempotency_conflict |
409 | This Idempotency-Key is still in progress, or its answer was streamed or not stored (zero retention). The message names the original request. |
Yes, if in progress | Wait and retry; otherwise inspect the original with GET /v1/requests/{id}. |
provider_credential_missing |
409 | A BYOK key, and your organization has no provider key for this model’s provider. | No | Add the provider key in the dashboard. |
duplicate_request |
409 | The request id was already used. Rare. | Yes | Retry the call. |
idempotency_key_reused |
422 | The same Idempotency-Key with a different body. |
No | Use a new key for a new request. |
429 — rate limits
Section titled “429 — rate limits”| Code | Meaning | Retry? | What to do |
|---|---|---|---|
rate_limited |
Requests or tokens per minute exceeded for this key, a platform rate limit, or the model’s provider limiting it. | Yes | Wait for Retry-After. See Rate limits. |
too_many_in_flight |
Too many calls in progress at once. | Yes | Wait for Retry-After, or lower your concurrency. |
5xx — the API could not serve
Section titled “5xx — the API could not serve”| Code | HTTP | Meaning | Retry? | What to do |
|---|---|---|---|---|
not_implemented |
501 | The key is in confidential mode, which is not available yet. |
No | Use a standard or zero_retention key. |
admission_unavailable |
503 | A check (key, catalog, spend control, idempotency) could not be completed, so the call was refused. | Yes | Retry after Retry-After. |
admission_unavailable |
503 | A BYOK key, and your organization’s BYOK pricing is not set up yet. The message says so. It still carries Retry-After: 60, so the official SDKs retry it on their own. |
No | Ask your administrator to complete the BYOK setup. Lower maxRetries, or stop on this message, so your code does not wait on retries that cannot succeed. |
pricing_unavailable |
503 | No price is configured for your organization yet. | No | Contact support. |
model_unavailable |
503 | The model is temporarily disabled or failing. | Yes | Retry after Retry-After, or fall back to another model. |
service_unavailable |
503 | The model could not be reached, or a service is temporarily down. | Yes | Retry after Retry-After. |
request_timeout |
504 | A non-streamed call ran out of time. The answer carries x-should-retry: false, so the official SDKs do not resend it. |
No | Retry it with stream: true. See Long calls. |
Long calls
Section titled “Long calls”A call runs for at most about five minutes (300 seconds). The API ends it itself just before that limit, so you get a clean answer rather than a dropped connection:
- Non-streamed —
504 request_timeout, withx-should-retry: false. A non-streamed answer sends nothing until it is complete, and its time allowance can be shorter than the five minutes. - Streamed — an in-band error with
code: "request_timeout"(see below).
For long reasoning or long tool runs, use stream: true: tokens arrive as they are produced, and the stream carries a : keep-alive comment every 15 seconds so proxies do not close an idle connection. SSE clients and the official SDKs ignore these comment lines.
Errors in a stream
Section titled “Errors in a stream”Once a streamed answer has started, the HTTP status is already 200. A failure after that point is reported in the stream, with the same codes as above:
-
Chat Completions — an error event, then the end of the stream. There is no final usage chunk.
data: {"error":{"message":"The model could not be reached","type":"api_error","code":"service_unavailable"}}data: [DONE] -
Responses — a
response.failedevent, whoseresponse.errorcarries thecodeandmessage.
A non-streamed call that fails once the model was called answers with the matching status instead — 400 invalid_request if the provider refused the request, 429 rate_limited, 503 service_unavailable, or 504 request_timeout. Check the call later with GET /v1/requests/{id} and the X-Request-Id you received.
What a failed call counts against your caps
Section titled “What a failed call counts against your caps”Every call reserves its worst-case cost on your key (and its parent) before it runs. When a Chat Completions or Responses call fails, the reservation is closed as follows:
| What happened | Counted against the key’s spend caps |
|---|---|
Refused before the model was called (any 4xx refusal above, or a check that could not complete) |
Nothing — the reservation is released. |
The provider refused the request before doing any work (its 400, 401, 403, 404, 413, 422 or 429), with nothing produced |
Nothing — the reservation is released. |
| The provider reported usage | The real cost, at your organization’s price. |
| A provider server error, or a lost connection, with nothing produced and no usage reported | The input share of the reserved estimate (at least one token). |
Anything else with no usage reported — output already streamed, the API’s own time limit (504), another provider status |
The full reserved estimate. |
When the provider reported usage, the call is billed for it like any other call. Otherwise, the amount settled for a failed call counts only against the key’s spend caps (and its parent’s) — it is not billed: no credits are debited, and it adds nothing to the cost in /v1/usage or /v1/credits. You see it as cost.settled_usd on GET /v1/requests/{id} and in the key’s spend on GET /v1/key.
On Embeddings, every failure releases the reservation: a failed embeddings call counts nothing.
A call refused by the API’s own checks before the model is called (validation, policy, caps, credits, rate limits) leaves no request record: GET /v1/requests/{id} answers 404 not_found for its X-Request-Id.
Handling errors in code
Section titled “Handling errors in code”curl -sS -D - https://api.belarel.com/v1/chat/completions \ -H "Authorization: Bearer $BELAREL_API_KEY" \ -H "Content-Type: application/json" \ -d '{"model": "<model-id>", "messages": [{"role": "user", "content": "Hello"}]}'# -D - prints the headers: look for X-Request-Id and, on 429/503, Retry-After.import osimport openaifrom openai import OpenAI
client = OpenAI(base_url="https://api.belarel.com/v1", api_key=os.environ["BELAREL_API_KEY"])
try: client.chat.completions.create( model="<model-id>", messages=[{"role": "user", "content": "Hello"}], )except openai.APIStatusError as e: body = e.response.json() error = body.get("error", {}) print(e.status_code, error.get("code"), error.get("param"), e.response.headers.get("x-request-id")) for decision in body.get("belarel", {}).get("policy", {}).get("decisions", []): print("policy:", decision["rule"], decision["verdict"])const res = await fetch('https://api.belarel.com/v1/chat/completions', { method: 'POST', headers: { Authorization: `Bearer ${process.env.BELAREL_API_KEY}`, 'Content-Type': 'application/json', }, body: JSON.stringify({ model: '<model-id>', messages: [{ role: 'user', content: 'Hello' }] }),});
if (!res.ok) { const body = await res.json(); console.error(res.status, body.error?.code, res.headers.get('x-request-id')); for (const d of body.belarel?.policy?.decisions ?? []) console.error('policy:', d.rule, d.verdict); const retryable = res.status === 429 || (res.status >= 500 && res.headers.get('x-should-retry') !== 'false'); if (retryable && res.headers.get('retry-after')) { const waitS = Number(res.headers.get('retry-after')); // retry after waitS seconds, with the same Idempotency-Key }}Contacting support
Section titled “Contacting support”Include:
- The
X-Request-Id(req_…) of the failing call — orbelarel.request_idfrom the body. - The HTTP status and
error.code. - The time of the call (UTC) and the key’s prefix (
bel_live_…), never the full key.

