Spend control patterns
Belarel checks money before a call runs, not after the bill arrives. This page puts the available controls together into patterns you can apply today: what to cap, how to test without spending, how to retry without paying twice, and how to find out when a limit is hit.
Know which refusal you hit
Section titled “Know which refusal you hit”Every spend refusal is a 4xx you can act on, and nothing is charged for it.
| Code | HTTP | Meaning | What to do |
|---|---|---|---|
key_cap_exceeded |
402 | The key’s daily or monthly USD cap — or its parent’s — is reached. The message says which. | Wait for the window to reset, or raise the cap. |
max_cost_per_request |
402 | This call’s estimated cost is above your organization’s per-request ceiling. | Lower max_tokens (or max_output_tokens on Responses). |
insufficient_credits |
402 | Your organization has no credits left for this call. | Top up, or contact your organization’s admin. |
budget_exceeded |
402 | An organization budget blocks the call. | Contact your organization’s admin. |
spend_refused |
402 | Your organization’s spend controls refused the call for another reason. | Contact your organization’s admin. |
max_tokens_exceeds_key_limit |
400 | max_tokens is above the key’s max_output_tokens. |
Ask for fewer tokens, or raise the key’s ceiling. |
rate_limited |
429 | Requests or tokens per minute exceeded. | Wait for Retry-After. |
too_many_in_flight |
429 | Too many concurrent calls on this key. | Wait for Retry-After, or lower concurrency. |
One more refusal is not a 4xx: 503 pricing_unavailable means no price is configured for your organization yet. It is not charged either, and a retry will not fix it — contact support.
Cap every key
Section titled “Cap every key”Each key carries its own limits: usd_day, usd_month, rpm, tpm, max_in_flight, and max_output_tokens. Set them when you create the key in the dashboard, or on sub-keys through POST /v1/keys and PATCH /v1/keys/{id}.
Before a call runs, Belarel reserves its worst-case cost: the prompt plus the maximum output tokens, at your organization’s price. If that reservation would push the key — or its parent — past a USD cap, the call is refused with 402 key_cap_exceeded. When the call finishes, the reservation is replaced by the real cost. A call the provider refuses before doing any work costs nothing; a Chat Completions or Responses call that fails without reporting usage counts the input part of its estimate, or the full estimate, against the key’s caps only — it is not billed and does not appear in your credits. See Admission & settlement.
Use separate keys for separate jobs — one per service, environment, or customer — so a runaway job only exhausts its own cap. For one key per end customer, see One sub-key per customer.
Use sandbox keys in CI
Section titled “Use sandbox keys in CI”A bel_test_… key is served by a deterministic sandbox instead of a real model. Its calls cost nothing: belarel.cost reads usd: 0 with billing: "sandbox". Use the same model ids as in production; the key’s rate and concurrency limits still apply.
# CI environmentBELAREL_API_KEY=bel_test_...Your test suite then exercises the real request path — authentication, tags, validation, error handling — without spending. See Sandbox.
Respect the per-request ceiling
Section titled “Respect the per-request ceiling”Your organization’s contract can set a maximum cost for a single request. Belarel compares the call’s worst-case estimate against it before running, and refuses with 402 max_cost_per_request if it is above, with param pointing at the token limit to lower. The ceiling is set by your contract, not by the API; it does not apply to sandbox keys.
Retry without paying twice
Section titled “Retry without paying twice”Send an Idempotency-Key header (1–255 characters) that identifies the logical operation, and reuse it on every retry. Keys are scoped to your API key and kept for 24 hours.
curl https://api.belarel.com/v1/chat/completions \ -H "Authorization: Bearer $BELAREL_API_KEY" \ -H "Content-Type: application/json" \ -H "Idempotency-Key: invoice-summary-8841" \ -d '{"model": "<model-id>", "messages": [{"role": "user", "content": "Summarize invoice 8841"}], "max_tokens": 200}'r = client.chat.completions.create( model="<model-id>", messages=[{"role": "user", "content": "Summarize invoice 8841"}], max_tokens=200, extra_headers={"Idempotency-Key": "invoice-summary-8841"},)const r = await client.chat.completions.create( { model: '<model-id>', messages: [{ role: 'user', content: 'Summarize invoice 8841' }], max_tokens: 200 }, { headers: { 'Idempotency-Key': 'invoice-summary-8841' } },);| Situation | Answer |
|---|---|
| Same key, same body, first call finished (non-streamed) | The stored answer, with Idempotent-Replayed: true. The model is not called again. |
| Same key while the first call is still running | 409 idempotency_conflict, naming the original request. |
| Same key, first call was streamed or used a zero-retention key | 409 idempotency_conflict — those answers are never stored. |
| Same key, different body | 422 idempotency_key_reused. |
| First call was refused or failed (non-streamed) | The key is freed; the retry runs normally. |
| First call was streamed and failed partway | The key stays used: 409 idempotency_conflict for 24 hours. Retry under a new key. |
A 504 request_timeout is the exception to blind retries: it carries x-should-retry: false, so the official SDKs do not replay it, because the same call would run out of time again. Retry it with stream: true.
See Idempotency.
Monitor spend
Section titled “Monitor spend”Three read-only endpoints, none of which cost anything:
- GET
/v1/key— any key reads itself:usage.dayandusage.monthwithspent_usd,reserved_usd,limit_usd,remaining_usd, andresets_at, plus its current rate-limit state. - GET
/v1/credits(usage:read) — your organization’sincluded,consumed, andbalance, in credits. - GET
/v1/usage(usage:read) — requests, tokens, credits, andcost_usdfor the key and its sub-keys, grouped bytotal,key,model,day, ortag:<name>, over up to 93 days.
const BASE = 'https://api.belarel.com/v1';const auth = { Authorization: `Bearer ${process.env.BELAREL_API_KEY}` };
const self = await (await fetch(`${BASE}/key`, { headers: auth })).json();const day = self.usage.day;if (day.limit_usd !== null && day.remaining_usd < 0.2 * day.limit_usd) { console.warn(`Key ${self.name}: ${day.remaining_usd} USD left today, resets ${day.resets_at}`);}
const qs = new URLSearchParams({ group_by: 'model' });const usage = await (await fetch(`${BASE}/usage?${qs}`, { headers: auth })).json();for (const row of usage.data) console.log(row.group, row.requests, row.cost_usd);Get notified when a cap is reached
Section titled “Get notified when a cap is reached”Subscribe to the api_key.cap_reached event in your organization’s dashboard (Webhooks tab). It fires when a call is refused because a key’s daily or monthly USD cap — or its parent’s — is reached, at most once per capped key per UTC day. The payload names the key (object.id, and its prefix in object.key, never the token) and carries data.key_name, data.caps, and data.spent (usd_day and usd_month, reservations included).
{ "id": "…", "type": "api_key.cap_reached", "version": "v1", "occurred_at": "…", "object": { "type": "api_key", "id": "…", "key": "bel_live_…", "saas_entity_id": "…" }, "data": { "organization_id": "…", "key_name": "customer acme", "caps": { "usd_day": 5, "usd_month": 50 }, "spent": { "usd_day": 4.98, "usd_month": 31.4 } }}object.saas_entity_id and data.organization_id are opaque account identifiers: store and compare them, but do not read meaning into them.
Verify the X-Belarel-Signature header before acting on a delivery; see Webhooks. More events will follow.

