Skip to content

Reasoning

Reasoning models think before they answer. You can ask for more or less of that thinking with a reasoning effort, read the reasoning the model chooses to show, and see how many tokens it spent. Reasoning tokens are output tokens: you pay for them.

Dialect Field
Chat Completions reasoning_effort
Responses reasoning.effort

Accepted values are minimal, low, medium and high. Any other value returns 400 invalid_request with param set to the field. If you send nothing, Belarel sets no effort and the model uses its provider’s default.

Terminal window
curl https://api.belarel.com/v1/chat/completions \
-H "Authorization: Bearer $BELAREL_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "openai/gpt-5", "reasoning_effort": "low",
"messages": [{"role": "user", "content": "Is 221 a prime number?"}]}'

Today the effort is applied to OpenAI models (openai/…). Other providers have no equivalent setting that Belarel translates.

Belarel does not drop the setting silently, and it does not refuse the call. The call runs at the model’s default, and belarel.policy.decisions says what happened:

"policy": {
"decisions": [
{ "rule": "org_api_catalog", "verdict": "allow", "detail": "residency any" },
{ "rule": "reasoning_effort", "verdict": "warn", "detail": "not applied: vendor \"anthropic\" takes no effort setting" }
]
}

When the effort is applied, the same rule appears with "verdict": "allow" and the effort as detail, for example "detail": "low".

Many models keep their reasoning private unless asked. Belarel asks OpenAI and Google models for a reasoning summary. Other models return reasoning only if their provider streams it by default. What you receive depends on the model.

Dialect Non-streamed Streamed
Chat Completions choices[0].message.reasoning_content (absent when there is none) delta.reasoning_content chunks
Responses a reasoning output item, with its text in summary[].text (type: "summary_text") response.reasoning_summary_text.delta events

You do not need to send reasoning back in the conversation. On the Responses API, reasoning items you include in input are accepted and skipped.

Reasoning tokens are counted inside the output tokens, and the details field tells you how many there were:

Dialect Output tokens Of which reasoning
Chat Completions usage.completion_tokens usage.completion_tokens_details.reasoning_tokens
Responses usage.output_tokens usage.output_tokens_details.reasoning_tokens

A high effort can spend many tokens before the first word of the answer. Leave room in max_tokens (or max_output_tokens): the budget covers reasoning and answer together.

If the budget runs out, the answer is cut short and the response says so: finish_reason is length on Chat Completions, and on Responses status is incomplete with incomplete_details.reason: "max_output_tokens". A reasoning model can spend the whole budget thinking and return little or no text.

Every call has a time limit of a little under 300 seconds. A non-streamed answer sends nothing until it is complete, so a long reasoning can reach that limit, or a proxy’s shorter idle timeout, before the first byte, and end in 504 request_timeout. With stream: true, reasoning and text reach you as they are generated, and Belarel sends a keep-alive comment every 15 seconds so the connection stays open. Use streaming for reasoning models with medium or high effort. See Streaming.