Skip to content

Changelog

Changes to the Belarel API, newest first. Additive changes and deprecations are both recorded here; see Versioning & deprecations for what each kind means for your integration.

Behavior changes

A set of corrections to the developer preview. Most make the API behave more like OpenAI’s; the ones marked Breaking can change what your integration receives.

  • Breaking — Chat Completions and Responses refuse any parameter they do not support, by name, with 400 unsupported_parameter and param set to the field. Before, unknown parameters were ignored. Refused by name: n above 1, logprobs, top_logprobs, logit_bias, store: true, prediction, audio, non-text modalities, a tool_choice that forces a tool, the legacy functions; on Responses also truncation: "auto" and background: true. See Unsupported parameters.
  • top_p, stop (up to 4), seed, presence_penalty, frequency_penalty and tool_choice: "none" are now forwarded to the model (top_p on Responses too). Values are validated: temperature from 0 to 2, top_p from 0 to 1, token budgets as positive integers.
  • New belarel.ignored_parameters on answers: the parameters the model’s provider ignored, when there are some. supported_parameters in GET /v1/models lists what is forwarded to each model.
  • Breaking — tools sent to a model that cannot call tools are refused with 400 model_lacks_tools. Before, they were removed and the call answered without tool calls.
  • finish_reason now reports why the model stopped: length when max_tokens cut the reply, content_filter, tool_calls, else stop. On Responses, a reply cut by max_output_tokens has status: "incomplete" and incomplete_details.reason: "max_output_tokens", and a stream ends with response.incomplete.
  • POST /v1/embeddings accepts encoding_format: "base64" (little-endian float32), the default of the official OpenAI SDKs, which now work with no setting.
  • Breaking — A missing key, or a bearer that is not a Belarel key, answers 401 invalid_api_key in the OpenAI shape on every endpoint, including /v1/chat/completions and /v1/responses.
  • When the model’s provider refuses a request as written, you get 400 invalid_request with the provider’s explanation, cleaned — fix the request rather than retrying it. A provider rate limit is a 429. On a BYOK key, a provider key the provider rejects is 400 byok_credential_rejected.
  • Breaking — Spend refusals use public codes only: insufficient_credits, budget_exceeded, spend_refused (402) and pricing_unavailable (503). price_book_missing is no longer returned.
  • Streamed failures carry the same code as non-streamed ones.
  • A call runs for at most about five minutes. A non-streamed call that runs out of time answers 504 request_timeout with x-should-retry: false; a streamed one ends with an in-band request_timeout error.
  • Streams send a : keep-alive comment every 15 seconds.
  • A call the provider refuses before doing any work costs nothing: its reservation is released.
  • A provider server error or a lost connection with nothing produced and no usage reported now counts the input share of the estimate against the key’s caps, instead of the full estimate. That amount counts against the caps only and is not billed. See what a failed call counts.
  • Embeddings whose price is not known when the answer is sent report cost.source: "pending".
  • Breaking — belarel.provider is removed from answers and from GET /v1/models; owned_by is the model’s vendor.
  • Breaking — GET /v1/requests/{id} no longer returns provider or gateway.
  • The keyless GET /v1/models shows the price list the platform publishes as its public price (pricing_basis: "public_price_book"), or no price when none is published.
  • The OpenAPI description lists bare hosts, so a client generated from it no longer requests /v1/v1/….
  • Required tags apply to model calls only: reading /v1/models, /v1/usage, /v1/credits, /v1/requests/{id} and managing /v1/keys no longer need them.
  • The idempotency fingerprint is now a keyed hash held by the server, for every key. A request sent with an Idempotency-Key before this change and repeated within its 24 hours answers 422 idempotency_key_reused.
Developer preview

The first public release of the API, under /v1. This entry describes the surface at preview launch; the entry above lists what changed since.

  • POST /v1/chat/completions — the OpenAI Chat Completions format, as JSON or server-sent events. Streams end with a chunk carrying usage and belarel, then data: [DONE]. Supports max_tokens / max_completion_tokens, temperature, image inputs on vision models, client-executed function tools (relayed as tool_calls), response_format (text, json_object, json_schema), reasoning_effort (minimal, low, medium, high) and web_search where your organization allows it.
  • POST /v1/responses — the OpenAI Responses format, stateless. Typed streaming events from response.created to response.completed. Function tools, web search, text.format and reasoning.effort. store: true, previous_response_id, item_reference and input_file are refused with 400.
  • POST /v1/embeddings — openai/text-embedding-3-small, 1536 dimensions, float encoding, up to 256 inputs of up to 8,000 characters each.
  • GET /v1/models — without a key, the public catalog at the platform’s starting price; with a key, the models that key may call at your organization’s price.
  • Each model carries catalog fields (name, description, context_length, architecture, top_provider, supported_parameters) and a belarel block: category, capabilities, residency zone, zero-retention and BYOK availability, and pricing in USD and credits.
  • Live (bel_live_…) and sandbox (bel_test_…) keys, created in the organization dashboard.
  • Scopes inference, embeddings, usage:read and keys:manage.
  • Per-key daily and monthly spend caps, requests and tokens per minute, calls in flight, maximum output tokens, allowed models, and required and default tags.
  • GET /v1/key — a key reads its own settings, spend, remaining cap and rate-limit state.
  • GET POST /v1/keys and PATCH DELETE /v1/keys/{id} — one level of sub-keys that never exceed their parent; reversible disable and definitive revoke.
  • Every call is reserved at your organization’s price before it runs, on the key and on its parent, then settled at its real cost or released when refused.
  • Every answer carries a belarel block: request id, model requested and served, provider (since removed), cost in USD and credits with its source and billing, data mode, and policy decisions.
  • Policy refusals name the rule in belarel.policy.decisions. A model your organization cannot see answers 404 model_not_found.
  • Errors use the OpenAI shape. 429 and 503 carry Retry-After.
  • X-RateLimit-Limit-Requests, X-RateLimit-Remaining-Requests and X-RateLimit-Reset-Requests on admitted answers.
  • X-Request-Id on the answers of the inference and read endpoints, errors included.
  • GET /v1/requests/{id} — one request’s status, model, provider (since removed), usage, cost and tags. Never its content.
  • GET /v1/usage — usage grouped by total, key, model, day or tag:<name>, over up to 93 days.
  • GET /v1/credits — your organization’s credit balance.
  • Idempotency-Key on inference calls, scoped to the key, kept for 24 hours.
  • Attribution tags from X-Belarel-Tags, the metadata body field and belarel_tags; X-Title and HTTP-Referer name your application.
  • Data modes per key: standard and zero_retention. confidential is coming soon and answers 501 not_implemented today.
  • Bring your own provider key (BYOK) for OpenAI and Anthropic models, configured in the organization dashboard.
  • Sandbox keys are served by a deterministic mock provider and never billed.
  • Webhook event api_key.cap_reached, signed with HMAC-SHA256 in X-Belarel-Signature. More events will follow.