Skip to content

Errors & debugging

The API speaks the OpenAI error dialect, so an official OpenAI SDK reads Belarel errors natively. This page lists the codes you can receive, what each means, and what to do about it.

Every error from a /v1 endpoint has this shape:

{
"error": {
"message": "Model \"<model-id>\" is not among the models this organization opened to its API keys",
"type": "permission_error",
"code": "model_not_allowed",
"param": "model"
},
"belarel": {
"request_id": "req_6f1c0e2a9b8d4c71a3e5f0b2d4c6e8a0",
"policy": {
"decisions": [{ "rule": "org_api_restriction", "verdict": "deny" }]
}
}
}
Field Meaning
error.code Stable, machine-readable code. Branch on this, never on message.
error.type The OpenAI error class, derived from the HTTP status (see below).
error.message A human-readable explanation. Wording can change.
error.param The request field at fault, when there is one — model, max_tokens, stop, Idempotency-Key… — else null.
belarel Present on policy refusals, on 402 max_cost_per_request and on 404 model_not_found: the request id and the policy rule that decided (unknown_model on a 404). SDKs ignore this extra key.

error.type follows the status: 400, 405, 422 → invalid_request_error · 401 → authentication_error · 402 → insufficient_quota · 403 → permission_error · 404 → not_found_error · 409 → conflict_error · 429 → rate_limit_error · 5xx → api_error.

  • A refusal is a 4xx. Money, policy, parameter and rate refusals all come back as 4xx, and they name what refused the call.
  • A 5xx means the API could not decide, could not reach the model, or ran out of time — so it refused. It never serves a call it could not check. Most 5xx answers are worth retrying after Retry-After. Four are not, because they answer the same until something changes: 501 not_implemented, 503 pricing_unavailable, a 503 admission_unavailable caused by an incomplete BYOK setup, and 504 request_timeout (retry it with stream: true instead).

When your organization’s policy refuses a model, the response names the rule in belarel.policy.decisions[], with verdict: "deny":

Rule Why the model was refused
contract_blocked_models Your contract blocks this model.
contract_allowed_models Your contract allows a list of models, and this one is not on it.
contract_allowed_providers The model is served by a provider your contract does not allow.
residency The model is not available in your organization’s residency zone.
org_api_restriction Your organization limited the models open to API keys, and this one is not among them.
org_api_disabled Your organization turned off this kind of model for API keys.
key_allowed_models This key is restricted to a list of models, and this one is not on it.
contract_max_cost_per_request The worst-case estimated cost of the call exceeds your contract’s maximum per request.

A model your organization cannot see at all is never a 403: it answers 404 model_not_found, exactly like a model that does not exist, with the generic decision {"rule": "unknown_model", "verdict": "deny"}.

Not every 403 model_not_allowed carries decisions: a refusal decided at the reservation step, or by a last availability check just before the model is called, comes without a belarel block.

Retry: Yes — retry after Retry-After (or with backoff). No — retrying the same request gives the same answer; change the request or the configuration first.

400 — the request cannot be served as sent

Section titled “400 — the request cannot be served as sent”
Code Meaning Retry? What to do
invalid_request A field is missing or invalid — a wrong type, temperature outside 0–2, top_p outside 0–1, a token budget that is not a positive integer, more than 4 stop sequences, a seed that is not an integer. Also used by /v1/embeddings, /v1/usage and /v1/keys bodies. param names the field. No Fix the field named in param.
invalid_request The model’s provider refused the request as written. The message starts with “The provider refused the request:” followed by the provider’s explanation, cleaned of links and secrets. No Fix the request; do not resend it unchanged.
unsupported_parameter A parameter this API does not support, named in param. See Unsupported parameters. No Remove the parameter.
validation_failed Chat Completions without model. No Send model.
invalid_messages messages is not a non-empty array of {role, content}, or a message uses a role other than system, user, assistant or tool — the developer role included. No Fix messages; send developer instructions as system.
invalid_input Responses input could not be read as a conversation. No Fix input.
invalid_tags A tag is malformed: bad key, value over 200 characters, more than 16 tags, or metadata is not an object of strings. No Fix the tags.
missing_tags A model call (chat, responses, embeddings) did not send a tag the key requires. The message lists them. Reads never need tags. No Add the required tags.
max_tokens_exceeds_key_limit The output budget is above this key’s limit. param is max_tokens on Chat Completions — even when you sent max_completion_tokens — and max_output_tokens on Responses. No Lower it, or omit it to use the key’s limit.
model_lacks_vision The request contains images and the model cannot read them. No Pick a vision-capable model.
model_lacks_tools The request declares tools, and the model cannot call tools. No Pick a model with the tools capability, or drop tools.
model_lacks_structured_output response_format / text.format asks for structured output the model cannot produce. param is always response_format, on Responses too. No Pick a model with the structured_output capability.
web_search_unavailable Web search was requested on a model with no native web search. No Pick another model, or drop web search.
connectors_require_user_session A non-empty connectors or hitl_approvals was sent; they are not available to API keys. No Remove them.
model_not_zdr A zero_retention key called a model that cannot be served with zero retention. No Pick a model with zero_retention_available: true.
byok_unavailable A BYOK key called a model (or embeddings) that cannot run on your own provider key. No Pick a model with byok_available: true.
byok_credential_rejected A BYOK key, and the provider refused your organization’s own provider key. No Check or replace the provider key in your organization’s BYOK settings.
model_not_allowed On /v1/embeddings: a model other than the one embedding model served. No Use the model listed for embeddings.
child_exceeds_parent A sub-key would exceed its parent (scopes, models, caps, rates, expiry). No Tighten the sub-key’s settings.
model_not_in_catalog A sub-key names models outside your organization’s API catalog. No Use models from GET /v1/models.
key_revoked You tried to disable or re-enable a revoked sub-key. No Create a new sub-key.

On Chat Completions and Responses, every top-level request field is either forwarded to the model, refused by name, or listed below as having no effect. A top-level field the endpoint does not know is refused with unsupported_parameter, like OpenAI does. So are these known fields:

  • Chat Completions: n above 1, logprobs: true, top_logprobs, a non-empty logit_bias, store: true, prediction, audio, modalities other than text, a tool_choice that forces a tool (only auto and none are accepted), and the legacy functions / function_call.
  • Responses: store: true, previous_response_id, item_reference and input_file inputs, tool types other than function and web_search, a tool_choice other than auto or none, truncation: "auto", background: true, top_logprobs.

Some fields are accepted and have no effect: user, stream_options, parallel_tool_calls and service_tier on Chat Completions; user, parallel_tool_calls and service_tier on Responses, where stream_options is refused with unsupported_parameter.

The check stops at the top level. Fields nested inside an accepted object — for example text.verbosity or reasoning.summary on Responses, or json_schema.strict inside response_format — are not checked: one the API does not use is ignored without an error. /v1/embeddings has no such check at all: it reads model, input, encoding_format and dimensions, and ignores any other field.

A forwarded parameter that the model’s provider then ignores (some ignore seed, the penalties, or top_p next to temperature) is not an error: the call succeeds and lists it in belarel.ignored_parameters. That list uses the Chat Completions names on both endpoints: on Responses, an ignored max_output_tokens is listed as max_tokens, and an ignored text.format as response_format.

Code Meaning Retry? What to do
invalid_api_key No key, a bearer that is not a Belarel key, or an unknown, revoked, expired or disabled key — deliberately one answer, on every endpoint. The one exception: GET /v1/models with no Authorization header answers 200 with the public catalog. No Send Authorization: Bearer bel_…, and check the key’s status in your dashboard.
Code Meaning Retry? What to do
key_cap_exceeded The key’s daily or monthly spend cap — or its parent’s — is reached. No Wait for the window to reset, or raise the cap. A webhook can alert you.
insufficient_credits Your organization does not have enough credits for this call. No Top up, or check your subscription.
budget_exceeded Your organization’s budget for this period is exhausted. No Ask an administrator to review the budget.
spend_refused Your organization’s spend controls refused the call. No Ask an administrator.
max_cost_per_request The worst-case estimated cost is above your contract’s per-request maximum. No Lower max_tokens.
Code Meaning Retry? What to do
scope_denied The key lacks the scope this endpoint needs. No Use a key with the scope.
model_not_allowed The model is not allowed. On a policy refusal, belarel.policy.decisions names the rule. No Pick another model, or change the policy.
model_not_priced The model has no price yet and cannot be served. No Pick another model.
web_search_not_allowed Web search is not available for this model in your organization. No Drop web search, or pick another model.
entitlement_denied Your organization’s plan does not allow this call. No Contact your administrator.
model_ambiguous The model id could not be resolved to a single model. Rare. No Contact support with the X-Request-Id.
Code HTTP Meaning Retry? What to do
model_not_found 404 The model does not exist or is not available to your organization. Carries the decision unknown_model. No List models with GET /v1/models.
not_found 404 No such request (/v1/requests/{id}) or sub-key (/v1/keys/{id}) for this key. No Check the id and the key you use.
method_not_allowed 405 Wrong HTTP method for this endpoint. No Use the documented method.
Code HTTP Meaning Retry? What to do
idempotency_conflict 409 This Idempotency-Key is still in progress, or its answer was streamed or not stored (zero retention). The message names the original request. Yes, if in progress Wait and retry; otherwise inspect the original with GET /v1/requests/{id}.
provider_credential_missing 409 A BYOK key, and your organization has no provider key for this model’s provider. No Add the provider key in the dashboard.
duplicate_request 409 The request id was already used. Rare. Yes Retry the call.
idempotency_key_reused 422 The same Idempotency-Key with a different body. No Use a new key for a new request.
Code Meaning Retry? What to do
rate_limited Requests or tokens per minute exceeded for this key, a platform rate limit, or the model’s provider limiting it. Yes Wait for Retry-After. See Rate limits.
too_many_in_flight Too many calls in progress at once. Yes Wait for Retry-After, or lower your concurrency.
Code HTTP Meaning Retry? What to do
not_implemented 501 The key is in confidential mode, which is not available yet. No Use a standard or zero_retention key.
admission_unavailable 503 A check (key, catalog, spend control, idempotency) could not be completed, so the call was refused. Yes Retry after Retry-After.
admission_unavailable 503 A BYOK key, and your organization’s BYOK pricing is not set up yet. The message says so. It still carries Retry-After: 60, so the official SDKs retry it on their own. No Ask your administrator to complete the BYOK setup. Lower maxRetries, or stop on this message, so your code does not wait on retries that cannot succeed.
pricing_unavailable 503 No price is configured for your organization yet. No Contact support.
model_unavailable 503 The model is temporarily disabled or failing. Yes Retry after Retry-After, or fall back to another model.
service_unavailable 503 The model could not be reached, or a service is temporarily down. Yes Retry after Retry-After.
request_timeout 504 A non-streamed call ran out of time. The answer carries x-should-retry: false, so the official SDKs do not resend it. No Retry it with stream: true. See Long calls.

A call runs for at most about five minutes (300 seconds). The API ends it itself just before that limit, so you get a clean answer rather than a dropped connection:

  • Non-streamed — 504 request_timeout, with x-should-retry: false. A non-streamed answer sends nothing until it is complete, and its time allowance can be shorter than the five minutes.
  • Streamed — an in-band error with code: "request_timeout" (see below).

For long reasoning or long tool runs, use stream: true: tokens arrive as they are produced, and the stream carries a : keep-alive comment every 15 seconds so proxies do not close an idle connection. SSE clients and the official SDKs ignore these comment lines.

Once a streamed answer has started, the HTTP status is already 200. A failure after that point is reported in the stream, with the same codes as above:

  • Chat Completions — an error event, then the end of the stream. There is no final usage chunk.

    data: {"error":{"message":"The model could not be reached","type":"api_error","code":"service_unavailable"}}
    data: [DONE]
  • Responses — a response.failed event, whose response.error carries the code and message.

A non-streamed call that fails once the model was called answers with the matching status instead — 400 invalid_request if the provider refused the request, 429 rate_limited, 503 service_unavailable, or 504 request_timeout. Check the call later with GET /v1/requests/{id} and the X-Request-Id you received.

What a failed call counts against your caps

Section titled “What a failed call counts against your caps”

Every call reserves its worst-case cost on your key (and its parent) before it runs. When a Chat Completions or Responses call fails, the reservation is closed as follows:

What happened Counted against the key’s spend caps
Refused before the model was called (any 4xx refusal above, or a check that could not complete) Nothing — the reservation is released.
The provider refused the request before doing any work (its 400, 401, 403, 404, 413, 422 or 429), with nothing produced Nothing — the reservation is released.
The provider reported usage The real cost, at your organization’s price.
A provider server error, or a lost connection, with nothing produced and no usage reported The input share of the reserved estimate (at least one token).
Anything else with no usage reported — output already streamed, the API’s own time limit (504), another provider status The full reserved estimate.

When the provider reported usage, the call is billed for it like any other call. Otherwise, the amount settled for a failed call counts only against the key’s spend caps (and its parent’s) — it is not billed: no credits are debited, and it adds nothing to the cost in /v1/usage or /v1/credits. You see it as cost.settled_usd on GET /v1/requests/{id} and in the key’s spend on GET /v1/key.

On Embeddings, every failure releases the reservation: a failed embeddings call counts nothing.

A call refused by the API’s own checks before the model is called (validation, policy, caps, credits, rate limits) leaves no request record: GET /v1/requests/{id} answers 404 not_found for its X-Request-Id.

Terminal window
curl -sS -D - https://api.belarel.com/v1/chat/completions \
-H "Authorization: Bearer $BELAREL_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "<model-id>", "messages": [{"role": "user", "content": "Hello"}]}'
# -D - prints the headers: look for X-Request-Id and, on 429/503, Retry-After.

Include:

  1. The X-Request-Id (req_…) of the failing call — or belarel.request_id from the body.
  2. The HTTP status and error.code.
  3. The time of the call (UTC) and the key’s prefix (bel_live_…), never the full key.