Skip to content

Overview

The Belarel API is one OpenAI-compatible door to the models your organization has enabled. You call it with the official OpenAI SDKs, and every call is governed on the way: admitted on its key, priced and reserved before it runs, checked against your organization’s policies and budget, then settled at what it really cost. Read this page first to learn the model the rest of the docs build on.

https://api.belarel.com/v1

Every endpoint lives under /v1. If you already use an OpenAI SDK, you only change two things: the base URL and the API key.

Each inference request goes through the same stages, in this order. Each stage can refuse the call, and every refusal comes back as a typed error you can act on.

Stage What is decided
Admitted The key is valid and active, has the scope the endpoint needs, is within its requests-per-minute and in-flight limits, and carries its required tags. The model is then checked against your organization’s catalog and the key’s own model list.
Priced and reserved The worst case of the call (prompt plus the maximum output) is priced at your organization’s price and reserved against the key’s daily and monthly spend caps, and its parent’s caps when it is a sub-key.
Checked Your contract’s ceiling per request, and your organization’s credits and budgets.
Served The model runs. You get the answer as JSON or as a stream.
Settled The reservation becomes the real cost. If the call is refused before the model is called, or the provider refuses it before doing any work, the reservation is released and nothing is spent. If the provider fails without reporting usage, the input share of the estimate, or the full estimate, counts against the key’s spend caps — it is not billed. See what a failed call counts.

When the door cannot reach a decision (for example, its catalog or ledger is briefly unavailable), it refuses with a 503 and a Retry-After header. It never serves a call it could not price or check.

Keys come in two modes, and the prefix tells you which one you hold:

Prefix Mode Served by Billed
bel_live_… Live The real model provider Yes, from your organization’s credits
bel_test_… Sandbox A deterministic mock provider Never

Start with a sandbox key: your code runs end to end, with the same request and response shapes, and no money moves. See Sandbox & testing.

These endpoints speak the OpenAI wire format, so official SDKs work unchanged:

  • POST /v1/chat/completions — JSON or server-sent events
  • POST /v1/responses — the Responses format, stateless
  • POST /v1/embeddings
  • GET /v1/models

An official OpenAI SDK works with no extra setting — embeddings included. Errors use the OpenAI shape { "error": { "message", "type", "code", "param" } } on every endpoint, so the SDKs raise their usual exception classes; a missing or invalid key is always 401 invalid_api_key. Each top-level Chat Completions and Responses parameter is forwarded to the model, accepted with a documented no-effect, or refused by name with 400 unsupported_parameter — never dropped in silence. Nested fields are not checked. See Authentication & keys and Errors & debugging.

On top of that, a successful answer from these four endpoints carries two things OpenAI does not send:

  • An X-Request-Id header (req_…), the identifier to log and to give support. Most refusals carry it too.
  • A belarel object next to the OpenAI fields. On an inference call it holds the request id, the model requested and the model that served the call, its cost in USD and credits, the data mode, the policy decisions that let it through, and — when there are some — the parameters the model’s provider ignored (ignored_parameters). On /v1/models it holds the request id and catalog information. SDKs ignore unknown fields, so it never breaks a client.

These have no OpenAI equivalent. Call them with any HTTP client.

Endpoint Scope Purpose
GET /v1/key none The calling key reads its own settings, spend and limits
GET /v1/credits usage:read Your organization’s credit balance
GET /v1/usage usage:read Usage grouped by key, model, day or tag
GET /v1/requests/{id} usage:read One request as the door saw it (never its content)
GET POST /v1/keys keys:manage List and create sub-keys
PATCH DELETE /v1/keys/{id} keys:manage Update, disable or revoke a sub-key