Admission & settlement
Every inference call goes through the same sequence: Belarel checks your key, checks the model against your organization’s policy, reserves the call’s worst-case price, serves it, then settles the reservation at the real cost. Knowing the order tells you which error you will see first, and why a spend cap can refuse a call that would have turned out cheap.
The lifecycle of one call
Section titled “The lifecycle of one call”This applies to POST /v1/chat/completions, POST /v1/responses and POST /v1/embeddings.
-
Request id. Belarel assigns a
req_…id before doing any work and returns it in theX-Request-Idheader, on refusals too. See Request inspection. -
Key admission. The key must exist, be active and unexpired, and carry the scope the endpoint needs. A sub-key is refused if its parent is disabled, revoked or expired. The call counts against the key’s requests-per-minute window, and the key’s calls in flight are counted. Tags are parsed here, so a missing required tag is refused before any money is involved. A key in
confidentialdata mode stops here with501 not_implemented— confidential inference is coming soon. -
Validation. The body is parsed in the endpoint’s format (Chat Completions, Responses or Embeddings). On Chat Completions and Responses, a top-level field the endpoint does not support is refused by name with
400 unsupported_parameterrather than ignored; nested fields are not checked. Embeddings ignores fields it does not read. See Errors. -
Model and policy. The model must be in your key’s catalog: exposed to the API, available to your organization, and allowed by your contract, your organization’s choices and the key itself. See Policies & catalog.
-
Capabilities and output budget. Images need a vision model, tools need a model that can call tools (
400 model_lacks_tools), web search must be allowed for the model, structured output must be supported.max_tokenscannot exceed the key’smax_output_tokens; when you omit it, the key’s limit becomes the ceiling. A BYOK key is checked against your organization’s provider key here. -
Idempotency. If you sent an
Idempotency-Key, a replay or a conflict is answered now, before any money. See Idempotency. -
Reservation. Belarel estimates the call’s worst-case price at your organization’s prices and reserves it on the key — and on its parent, for a sub-key — in one atomic step that compares it with the daily and monthly spend caps. The tokens-per-minute limit is checked here, then your contract’s maximum cost per request.
-
Organization credits and budgets. The estimate must fit your organization’s credit balance and any budget it has set. For a BYOK key too, this check — like the reservation and the maximum cost per request — uses the full-price estimate; only settlement is at the BYOK fee.
-
Serving. The model runs. Your answer — JSON or a stream — closes with
usageand thebelarelblock. -
Settlement. The reservation becomes the real spend, and the difference goes back to your caps.
The reservation
Section titled “The reservation”The reservation is what makes caps hold under concurrency. Each call reserves its estimate before it runs, and a cap compares spent + reserved + this estimate with its limit. Two parallel streams cannot both slip under the last dollar of a daily cap.
The estimate is a worst case: the size of your prompt plus the full output ceiling (max_tokens, or the key’s max_output_tokens when you omit it), repeated over the tool-call steps the call may take, at your organization’s sale price. A sandbox key (bel_test_…) reserves zero.
Settlement and release
Section titled “Settlement and release”When the provider reports usage for a call, the reservation settles at that cost — whether the call succeeded or failed — and the call is billed for it. The rules below apply to Chat Completions and Responses when it does not. On Embeddings, every failure releases the reservation: a failed embeddings call counts nothing.
| What happened | What Belarel does |
|---|---|
| The call was served | Settles at the real cost reported for the call. |
| The call was refused after its reservation but before the model was called (credits, budget, model switched off, provider circuit open, entitlement, organization rate limit) | Releases the reservation: nothing is spent. |
The provider refused the request before doing any work — it answered 400, 401, 403, 404, 413, 422 or 429, and nothing was produced |
Releases the reservation: nothing is spent. |
The provider failed with a server error (5xx), or the stream died without a status, before anything was produced and without reporting usage |
Settles at the input part of the estimate: the estimate × prompt tokens ÷ estimated tokens, counting the prompt as at least one token. |
Belarel’s own time limit stopped the call (504 request_timeout), the provider failed with another status (such as 402, 408 or 409), or output had already been produced before the failure — and no usage was reported |
Settles at the full estimate. |
| You disconnected in the middle of a stream | The model is stopped; the call is billed for what it consumed. |
| No cost could be computed for a served call | Settles at the estimate. A served call is never counted as free. |
| The reservation was never closed (the server stopped mid-call) | Once it is older than 15 minutes, settles it at the estimate on the key’s next request. |
“Produced” means a token, a reasoning delta or a tool call reached your answer. For a failed call, an amount settled at the estimate or at its input part counts only against the key’s spend caps (and its parent’s, for a sub-key) — it is not billed: no credits are debited, and it adds nothing to the cost in /v1/usage or /v1/credits. You see it as cost.settled_usd on GET /v1/requests/{id} and in the key’s spend on GET /v1/key.
Released and settled reservations both stop counting as “in flight” for the key.
When the cost is not known yet
Section titled “When the cost is not known yet”After the model finishes, Belarel waits up to 2 seconds for the call’s cost. If the figure is not ready, the answer still goes out, with:
"cost": { "usd": null, "credits": null, "source": "pending", "billing": "platform" }Settlement continues in the background. Read the final figure later with GET /v1/requests/{id}, or in your aggregates with GET /v1/usage. See Cost & usage.
Fail closed
Section titled “Fail closed”If Belarel cannot reach a decision — the key check, the catalog, the reservation, the credit check or a governance check is unavailable — it refuses the call rather than serving it unchecked. A refusal of your request is a 4xx — including a provider that refuses the request as written (400 invalid_request, “The provider refused the request: …”), which you fix rather than retry. A 5xx means Belarel could not decide, could not reach the model, or ran out of time, never that your request was wrong. Three 5xx answers will not be fixed by a plain retry: 501 not_implemented (confidential inference, coming soon), 503 pricing_unavailable (no price is configured for your organization yet — contact support) and 503 admission_unavailable when your organization’s BYOK setup is not complete — see Errors.
| Code | HTTP | Meaning | What to do |
|---|---|---|---|
admission_unavailable |
503 | A check could not be completed. | Retry after Retry-After (5 s). |
model_unavailable |
503 | The model is temporarily switched off or its provider is failing. | Retry after Retry-After when present, or pick another model. |
service_unavailable |
503 | The model could not be reached. | Retry after Retry-After (5 s). Each such 503 may count the input part of its estimate against your key’s caps. |
request_timeout |
504 | The call exceeded the time allowed. It carries x-should-retry: false, so the official SDKs do not replay it. |
Retry with stream: true. |

