Skip to content

Chat Completions

POST /v1/chat/completions speaks the OpenAI Chat Completions dialect. Point an official OpenAI SDK at https://api.belarel.com/v1 and your existing code works. Every call is admitted on your key, checked against your organization’s policies and budget, and settled at its real cost. The response tells you what it cost.

The key needs the inference scope.

Terminal window
curl https://api.belarel.com/v1/chat/completions \
-H "Authorization: Bearer $BELAREL_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "openai/gpt-5",
"messages": [
{"role": "system", "content": "You answer in one sentence."},
{"role": "user", "content": "What is an idempotency key?"}
],
"max_tokens": 200
}'

Replace openai/gpt-5 with a model your key may call. GET /v1/models lists them. See Models.

Every top-level field you send is either forwarded to the model, accepted with no effect, or refused by name. A field Belarel does not know is refused too, as OpenAI does: you get 400 unsupported_parameter with param set to that field, and the call costs nothing. If several fields are at fault, the error names the first one. The check covers top-level fields only: a field nested inside an accepted object (for example json_schema.strict in response_format) is not checked, and one Belarel does not use is ignored.

Field Notes
model Required. A model id from GET /v1/models. A model your organization cannot see returns 404 model_not_found.
messages Required. A non-empty array. Roles: system, user, assistant, tool. content is a string or an array of parts.
stream true streams server-sent events. Omitted or false returns one JSON object. See Streaming.
max_tokens, max_completion_tokens The output budget, a positive integer. If both are present, max_completion_tokens wins. If neither is present, your key’s output ceiling applies. A value above that ceiling returns 400 max_tokens_exceeds_key_limit, with param: "max_tokens" even when you sent max_completion_tokens.
tools Function tools that you execute. See Tool calling.
response_format text, json_object or json_schema. See Structured outputs.
reasoning_effort minimal, low, medium or high. See Reasoning.
web_search Belarel extension. true turns on the provider’s native web search for this turn. See Tool calling.
metadata, belarel_tags Attribution tags (string values). See Tags and attribution.

These sampling parameters are validated, then passed to the model. A value of the wrong type or out of range returns 400 invalid_request with param naming the field (a string such as "0.7" is refused, not converted).

Field Accepted values
temperature A number from 0 to 2.
top_p A number from 0 to 1.
stop A non-empty string, or an array of at most 4 non-empty strings.
seed An integer.
presence_penalty, frequency_penalty A number from -2 to 2.
tool_choice "auto" or "none". With "none", your tools stay declared (a history with tool calls may need them) and the model does not call them.

Forwarded does not mean honored. Some providers drop some of these settings: for example, a provider may ignore seed or the penalties, or ignore top_p when temperature is also set. When that happens, the call still succeeds, and belarel.ignored_parameters lists what the provider dropped, by its OpenAI name:

"belarel": { "request_id": "req_0f1e…", "ignored_parameters": ["frequency_penalty", "seed"], "…": "…" }

The field is absent when nothing was dropped. Each model’s supported_parameters in GET /v1/models lists the parameters Belarel forwards to it.

user, stream_options, parallel_tool_calls, service_tier and chat_id are accepted so that existing code keeps working, but they change nothing on an API key. A streamed answer always ends with a usage chunk, whatever stream_options says. n: 1, logprobs: false and store: false are accepted for the same reason.

You send Result
n other than 1 400 unsupported_parameter. You always get one choice.
logprobs: true, or top_logprobs 400 unsupported_parameter
A non-empty logit_bias 400 unsupported_parameter
store: true 400 unsupported_parameter. Belarel stores no completions.
prediction 400 unsupported_parameter
audio, or modalities other than text 400 unsupported_parameter. Output is text only.
functions, function_call 400 unsupported_parameter. Use tools.
tool_choice that forces a tool ("required", or a named function) 400 unsupported_parameter. Only "auto" and "none" are supported.
A non-empty connectors or hitl_approvals 400 connectors_require_user_session. These are used by the Belarel web app, not by API keys. An empty array is accepted and has no effect.
Any other field, such as web_search_options or verbosity 400 unsupported_parameter, param set to the field.

The error message names the field, for example The "web_search_options" parameter is not supported by this API.

  • Text parts ({"type": "text", "text": "…"}) are joined into one string.
  • Image parts ({"type": "image_url", "image_url": {"url": "…"}}) are sent to the model on user messages. The URL can be https://… or a data: URL. The model must be vision-capable (belarel.capabilities.vision in GET /v1/models). If it is not, you get 400 model_lacks_vision.
  • Other part types are reduced to their text, or dropped if they have none. Images on non-user messages are dropped the same way.
  • Empty messages (text that is empty after flattening, with no tool calls) are skipped.

A non-streamed answer is a standard chat.completion object with Belarel additions:

{
"id": "chatcmpl-3f1c…",
"object": "chat.completion",
"created": 1759300000,
"model": "openai/gpt-5",
"choices": [
{
"index": 0,
"message": { "role": "assistant", "content": "A unique value you send so a retried request runs only once." },
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 31,
"completion_tokens": 14,
"total_tokens": 45,
"prompt_tokens_details": { "cached_tokens": 0 },
"completion_tokens_details": { "reasoning_tokens": 0 }
},
"belarel": {
"request_id": "req_0f1e…",
"model_requested": "openai/gpt-5",
"model_served": "openai/gpt-5",
"cost": { "usd": 0.00004, "credits": 0.04, "source": "estimated", "billing": "platform" },
"policy": { "decisions": [{ "rule": "org_api_catalog", "verdict": "allow", "detail": "residency any" }] },
"data_mode": "standard"
}
}
  • message.reasoning_content is present when the model returned visible reasoning. See Reasoning.

  • message.tool_calls is present when the model called one of your tools.

  • finish_reason says why the model stopped:

    Value Meaning
    stop The model finished its answer, or reached one of your stop sequences.
    length The output budget (max_tokens or your key’s ceiling) cut the answer short.
    content_filter The provider filtered the answer.
    tool_calls The turn ended on a call to one of your tools.
  • id identifies the completion. The request id is belarel.request_id. It is also in the X-Request-Id header of every answer from this endpoint, errors included (except 405).

Field Meaning
request_id The request id (req_…). Use it with GET /v1/requests/{id} and when you contact support.
model_requested, model_served The model you asked for, and the model that answered.
cost.usd, cost.credits What the call cost, at your organization’s prices.
cost.source metered, estimated, or pending when the cost was not known yet. A pending cost is reconciled later. Read it with request inspection.
cost.billing platform, byok (the amount is the platform fee, see BYOK), or sandbox (always 0).
policy.decisions The rules that let the call through, each with a verdict of allow or warn.
data_mode standard or zero_retention. See Data modes.
ignored_parameters Present only when the provider dropped some forwarded parameters. Lists them by their OpenAI names. See Forwarded to the model.

The official SDKs keep unknown fields, so you can read belarel without a custom client. In Python, use getattr(completion, "belarel", None).

Errors use the OpenAI shape {"error": {"message", "type", "code", "param"}}, so the SDKs raise their usual exception classes. A refusal by policy also carries belarel.policy.decisions, which names the rule that refused. The codes you are most likely to meet on this endpoint:

Code HTTP Meaning What to do
unsupported_parameter 400 A field this API does not support. param names it. Remove the field.
invalid_request 400 A field has a wrong type or value, or the provider refused the request (the message starts with The provider refused the request:). Fix the request. Retrying it unchanged will fail again.
model_lacks_tools 400 You sent tools to a model that cannot call tools. Pick a model with belarel.capabilities.tools: true.
request_timeout 504 A non-streamed call ran out of time. Retry with stream: true. See Streaming.

The full list is in Errors.