Skip to content

Authentication & keys

Every request is authenticated with an API key that belongs to your organization. The key is more than a credential: it carries the scopes, limits and rules that govern every call made with it. Read this page before you hand keys to services or to your own customers.

Authorization: Bearer bel_live_...

There is no other way to authenticate. A missing header, a bearer that is not a Belarel key, and a bel_… key that is unknown, revoked, expired or disabled all answer the same 401 invalid_api_key, in the OpenAI error shape, on every endpoint. The only request you can make without a key is the public model catalog, GET /v1/models with no Authorization header at all (see Models & catalog).

A token looks like bel_live_ab12cd34_…. The first part, up to the second underscore, is the key’s prefix: it identifies the key in the dashboard and in API answers without revealing it.

Prefix Mode Behavior
bel_live_ live Served by real providers, billed from your organization’s credits.
bel_test_ sandbox Served by a deterministic mock provider, never billed. See Sandbox & testing.

Each key also has an environment label — production, staging, development or test — for your own bookkeeping. It does not change how calls are served; the mode does.

Owners and administrators of your organization create keys in the organization dashboard, under Governance → API keys. At creation you choose the name, mode, environment, daily and monthly spend caps, requests per minute, maximum output tokens per request, zero data retention, billing on your own provider keys (BYOK), and whether the key can create sub-keys. Keys created there get the scopes inference, embeddings and usage:read, plus keys:manage when they can create sub-keys.

A key with keys:manage can then create sub-keys through the API, with finer settings: an allowed-model list, required and default tags, tokens per minute, calls in flight and an expiry.

Scope Grants
inference POST /v1/chat/completions, POST /v1/responses, and the keyed GET /v1/models
embeddings POST /v1/embeddings (and embedding models in your model list)
usage:read GET /v1/usage, /v1/credits, /v1/requests/{id}
keys:manage GET POST /v1/keys, PATCH DELETE /v1/keys/{id}
none GET /v1/key — any valid key can read itself

A call outside the key’s scopes answers 403 scope_denied. Listing models with a key also needs inference: an embeddings-only key gets 403 scope_denied from GET /v1/models. List models with a key that has inference, or call the endpoint without a key for the public catalog.

Setting Default When it is exceeded
Daily spend cap (USD) none 402 key_cap_exceeded
Monthly spend cap (USD) none 402 key_cap_exceeded
Requests per minute 60 429 rate_limited, with Retry-After
Tokens per minute none 429 rate_limited, with Retry-After
Calls in flight 20 429 too_many_in_flight, with Retry-After
Max output tokens per request 4096 400 max_tokens_exceeds_key_limit

Spend caps are checked before the call runs, against a reservation of its worst case: your prompt plus the maximum output, at your organization’s price. A call can therefore be refused while the cap still has some room. Lowering max_tokens lowers the reservation. When you omit max_tokens, the key’s maximum is used. Spend windows reset at midnight UTC and on the first day of the month, UTC.

A key can carry a list of model ids. A call to any other model answers 403 model_not_allowed, and belarel.policy.decisions names the rule key_allowed_models. The list narrows what your organization already allows; it never widens it. Without a list, the key can call every model of your organization’s catalog.

Required tags must be present on every model call — POST /v1/chat/completions, /v1/responses and /v1/embeddings — or the call answers 400 missing_tags before any money is reserved. Reads and key management (/v1/models, /v1/usage, /v1/credits, /v1/requests/{id}, /v1/key, /v1/keys) never need them: tags attribute model calls, they do not grant access. Default tags are added to every call; a tag you send yourself wins. Tags can come from the X-Belarel-Tags header, the metadata body field or belarel_tags. See Tags & attribution.

A key’s data mode decides how the content of its calls is handled:

  • standard — the default.
  • zero_retention — content is kept nowhere, and only providers with a zero-retention, no-training agreement may serve the key. A model that has none answers 400 model_not_zdr.
  • confidential — Coming soon. Not available yet; such a key answers 501 not_implemented.

See Data modes.

A key with the keys:manage scope creates sub-keys, for example one per customer of yours, each with its own caps and tags. Sub-keys follow strict rules:

  • One level. A sub-key cannot create sub-keys, and cannot hold keys:manage. Its scopes are a subset of inference, embeddings and usage:read.
  • Never above its parent. Scopes, allowed models, spend caps, rate limits, max output tokens and expiry must all fit inside the parent’s, or the request answers 400 child_exceeds_parent. A setting you leave out is inherited from the parent, never unlimited.
  • Inherited. Mode, environment and BYOK come from the parent. Data mode can only get stricter: a sub-key of a zero-retention key is zero-retention. The parent’s required tags stay required.
  • Counted twice. A sub-key’s spend is reserved against its own caps and its parent’s.
  • Allowed models must come from your organization’s catalog, or 400 model_not_in_catalog.
Terminal window
curl https://api.belarel.com/v1/keys \
-H "Authorization: Bearer $BELAREL_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"name": "customer-acme",
"scopes": ["inference"],
"caps": { "usd_day": 5, "rpm": 30 },
"default_tags": { "customer": "acme" }
}'

The answer (201) includes the sub-key’s token, shown this one time. Disable it with PATCH /v1/keys/{id} and {"disabled": true} (reversible), or revoke it with DELETE /v1/keys/{id} (definitive). Calls already in progress finish and are billed. See One sub-key per customer.

GET /v1/key returns the calling key’s settings, what it spent and reserved today and this month, what remains, when each window resets, and its current rate-limit state. It never returns the token. Any valid key can call it, even one at its in-flight limit.

Terminal window
curl https://api.belarel.com/v1/key \
-H "Authorization: Bearer $BELAREL_API_KEY"

The answer has the key’s public fields (id, name, prefix, status, mode, environment, data_mode, scopes, allowed_model_keys, caps, required_tags, default_tags, expires_at, last_used_at…), plus usage.day and usage.month (spent_usd, reserved_usd, limit_usd, remaining_usd, resets_at), rate_limit (requests_per_minute, requests_this_minute, tokens_per_minute, max_in_flight, in_flight) and sub_key_count.

An unknown, revoked, expired or disabled key — or a sub-key whose parent is in one of those states — always gets the same answer:

{
"error": {
"message": "Invalid, revoked, or expired API key",
"type": "authentication_error",
"code": "invalid_api_key",
"param": null
}
}

The answer never tells which case applies, so a key cannot be probed. Check the key’s status in the dashboard.

  • Store the token when it is shown. It is displayed once; Belarel keeps only a fingerprint.
  • Keep keys server-side. Never ship a key in browser, mobile or desktop code.
  • One key per service and environment, so you can cap, inspect and revoke each one alone.
  • Rotate from the dashboard. Rotate issues a new key with the same settings; the old token keeps working for up to 24 hours, then expires. Its sub-keys move under the new key. A sub-key is not rotated: revoke it and create a new one. The new key starts with fresh daily and monthly spend counters, while the old one keeps its own until it expires — so during that window the two together can spend up to about twice the daily cap. Retire the old token as soon as your services use the new one.
  • Revoke at once if a key leaks. New calls are refused immediately; calls in progress finish and are billed. Revoking a key also revokes its sub-keys.