Chat Completions
POST /v1/chat/completions speaks the OpenAI Chat Completions dialect. Point an official OpenAI SDK at https://api.belarel.com/v1 and your existing code works. Every call is admitted on your key, checked against your organization’s policies and budget, and settled at its real cost. The response tells you what it cost.
The key needs the inference scope.
A first call
Section titled “A first call”curl https://api.belarel.com/v1/chat/completions \ -H "Authorization: Bearer $BELAREL_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "openai/gpt-5", "messages": [ {"role": "system", "content": "You answer in one sentence."}, {"role": "user", "content": "What is an idempotency key?"} ], "max_tokens": 200 }'import osfrom openai import OpenAI
client = OpenAI(base_url="https://api.belarel.com/v1", api_key=os.environ["BELAREL_API_KEY"])
completion = client.chat.completions.create( model="openai/gpt-5", messages=[ {"role": "system", "content": "You answer in one sentence."}, {"role": "user", "content": "What is an idempotency key?"}, ], max_tokens=200,)print(completion.choices[0].message.content)print(getattr(completion, "belarel", None)) # cost, request id, policy decisionsimport OpenAI from 'openai';
const client = new OpenAI({ baseURL: 'https://api.belarel.com/v1', apiKey: process.env.BELAREL_API_KEY });
const completion = await client.chat.completions.create({ model: 'openai/gpt-5', messages: [ { role: 'system', content: 'You answer in one sentence.' }, { role: 'user', content: 'What is an idempotency key?' }, ], max_tokens: 200,});console.log(completion.choices[0].message.content);console.log((completion as unknown as { belarel: unknown }).belarel);Replace openai/gpt-5 with a model your key may call. GET /v1/models lists them. See Models.
Request fields
Section titled “Request fields”Every top-level field you send is either forwarded to the model, accepted with no effect, or refused by name. A field Belarel does not know is refused too, as OpenAI does: you get 400 unsupported_parameter with param set to that field, and the call costs nothing. If several fields are at fault, the error names the first one. The check covers top-level fields only: a field nested inside an accepted object (for example json_schema.strict in response_format) is not checked, and one Belarel does not use is ignored.
Read by Belarel
Section titled “Read by Belarel”| Field | Notes |
|---|---|
model |
Required. A model id from GET /v1/models. A model your organization cannot see returns 404 model_not_found. |
messages |
Required. A non-empty array. Roles: system, user, assistant, tool. content is a string or an array of parts. |
stream |
true streams server-sent events. Omitted or false returns one JSON object. See Streaming. |
max_tokens, max_completion_tokens |
The output budget, a positive integer. If both are present, max_completion_tokens wins. If neither is present, your key’s output ceiling applies. A value above that ceiling returns 400 max_tokens_exceeds_key_limit, with param: "max_tokens" even when you sent max_completion_tokens. |
tools |
Function tools that you execute. See Tool calling. |
response_format |
text, json_object or json_schema. See Structured outputs. |
reasoning_effort |
minimal, low, medium or high. See Reasoning. |
web_search |
Belarel extension. true turns on the provider’s native web search for this turn. See Tool calling. |
metadata, belarel_tags |
Attribution tags (string values). See Tags and attribution. |
Forwarded to the model
Section titled “Forwarded to the model”These sampling parameters are validated, then passed to the model. A value of the wrong type or out of range returns 400 invalid_request with param naming the field (a string such as "0.7" is refused, not converted).
| Field | Accepted values |
|---|---|
temperature |
A number from 0 to 2. |
top_p |
A number from 0 to 1. |
stop |
A non-empty string, or an array of at most 4 non-empty strings. |
seed |
An integer. |
presence_penalty, frequency_penalty |
A number from -2 to 2. |
tool_choice |
"auto" or "none". With "none", your tools stay declared (a history with tool calls may need them) and the model does not call them. |
Forwarded does not mean honored. Some providers drop some of these settings: for example, a provider may ignore seed or the penalties, or ignore top_p when temperature is also set. When that happens, the call still succeeds, and belarel.ignored_parameters lists what the provider dropped, by its OpenAI name:
"belarel": { "request_id": "req_0f1e…", "ignored_parameters": ["frequency_penalty", "seed"], "…": "…" }The field is absent when nothing was dropped. Each model’s supported_parameters in GET /v1/models lists the parameters Belarel forwards to it.
Accepted with no effect
Section titled “Accepted with no effect”user, stream_options, parallel_tool_calls, service_tier and chat_id are accepted so that existing code keeps working, but they change nothing on an API key. A streamed answer always ends with a usage chunk, whatever stream_options says. n: 1, logprobs: false and store: false are accepted for the same reason.
Refused by name
Section titled “Refused by name”| You send | Result |
|---|---|
n other than 1 |
400 unsupported_parameter. You always get one choice. |
logprobs: true, or top_logprobs |
400 unsupported_parameter |
A non-empty logit_bias |
400 unsupported_parameter |
store: true |
400 unsupported_parameter. Belarel stores no completions. |
prediction |
400 unsupported_parameter |
audio, or modalities other than text |
400 unsupported_parameter. Output is text only. |
functions, function_call |
400 unsupported_parameter. Use tools. |
tool_choice that forces a tool ("required", or a named function) |
400 unsupported_parameter. Only "auto" and "none" are supported. |
A non-empty connectors or hitl_approvals |
400 connectors_require_user_session. These are used by the Belarel web app, not by API keys. An empty array is accepted and has no effect. |
Any other field, such as web_search_options or verbosity |
400 unsupported_parameter, param set to the field. |
The error message names the field, for example The "web_search_options" parameter is not supported by this API.
Message content
Section titled “Message content”- Text parts (
{"type": "text", "text": "…"}) are joined into one string. - Image parts (
{"type": "image_url", "image_url": {"url": "…"}}) are sent to the model onusermessages. The URL can behttps://…or adata:URL. The model must be vision-capable (belarel.capabilities.visioninGET /v1/models). If it is not, you get400 model_lacks_vision. - Other part types are reduced to their
text, or dropped if they have none. Images on non-usermessages are dropped the same way. - Empty messages (text that is empty after flattening, with no tool calls) are skipped.
The response
Section titled “The response”A non-streamed answer is a standard chat.completion object with Belarel additions:
{ "id": "chatcmpl-3f1c…", "object": "chat.completion", "created": 1759300000, "model": "openai/gpt-5", "choices": [ { "index": 0, "message": { "role": "assistant", "content": "A unique value you send so a retried request runs only once." }, "finish_reason": "stop" } ], "usage": { "prompt_tokens": 31, "completion_tokens": 14, "total_tokens": 45, "prompt_tokens_details": { "cached_tokens": 0 }, "completion_tokens_details": { "reasoning_tokens": 0 } }, "belarel": { "request_id": "req_0f1e…", "model_requested": "openai/gpt-5", "model_served": "openai/gpt-5", "cost": { "usd": 0.00004, "credits": 0.04, "source": "estimated", "billing": "platform" }, "policy": { "decisions": [{ "rule": "org_api_catalog", "verdict": "allow", "detail": "residency any" }] }, "data_mode": "standard" }}-
message.reasoning_contentis present when the model returned visible reasoning. See Reasoning. -
message.tool_callsis present when the model called one of your tools. -
finish_reasonsays why the model stopped:Value Meaning stopThe model finished its answer, or reached one of your stopsequences.lengthThe output budget ( max_tokensor your key’s ceiling) cut the answer short.content_filterThe provider filtered the answer. tool_callsThe turn ended on a call to one of your tools. -
ididentifies the completion. The request id isbelarel.request_id. It is also in theX-Request-Idheader of every answer from this endpoint, errors included (except405).
The belarel block
Section titled “The belarel block”| Field | Meaning |
|---|---|
request_id |
The request id (req_…). Use it with GET /v1/requests/{id} and when you contact support. |
model_requested, model_served |
The model you asked for, and the model that answered. |
cost.usd, cost.credits |
What the call cost, at your organization’s prices. |
cost.source |
metered, estimated, or pending when the cost was not known yet. A pending cost is reconciled later. Read it with request inspection. |
cost.billing |
platform, byok (the amount is the platform fee, see BYOK), or sandbox (always 0). |
policy.decisions |
The rules that let the call through, each with a verdict of allow or warn. |
data_mode |
standard or zero_retention. See Data modes. |
ignored_parameters |
Present only when the provider dropped some forwarded parameters. Lists them by their OpenAI names. See Forwarded to the model. |
The official SDKs keep unknown fields, so you can read belarel without a custom client. In Python, use getattr(completion, "belarel", None).
Errors
Section titled “Errors”Errors use the OpenAI shape {"error": {"message", "type", "code", "param"}}, so the SDKs raise their usual exception classes. A refusal by policy also carries belarel.policy.decisions, which names the rule that refused. The codes you are most likely to meet on this endpoint:
| Code | HTTP | Meaning | What to do |
|---|---|---|---|
unsupported_parameter |
400 | A field this API does not support. param names it. |
Remove the field. |
invalid_request |
400 | A field has a wrong type or value, or the provider refused the request (the message starts with The provider refused the request:). |
Fix the request. Retrying it unchanged will fail again. |
model_lacks_tools |
400 | You sent tools to a model that cannot call tools. |
Pick a model with belarel.capabilities.tools: true. |
request_timeout |
504 | A non-streamed call ran out of time. | Retry with stream: true. See Streaming. |
The full list is in Errors.

