Reasoning
Reasoning models think before they answer. You can ask for more or less of that thinking with a reasoning effort, read the reasoning the model chooses to show, and see how many tokens it spent. Reasoning tokens are output tokens: you pay for them.
Set the effort
Section titled “Set the effort”| Dialect | Field |
|---|---|
| Chat Completions | reasoning_effort |
| Responses | reasoning.effort |
Accepted values are minimal, low, medium and high. Any other value returns 400 invalid_request with param set to the field. If you send nothing, Belarel sets no effort and the model uses its provider’s default.
curl https://api.belarel.com/v1/chat/completions \ -H "Authorization: Bearer $BELAREL_API_KEY" \ -H "Content-Type: application/json" \ -d '{"model": "openai/gpt-5", "reasoning_effort": "low", "messages": [{"role": "user", "content": "Is 221 a prime number?"}]}'import osfrom openai import OpenAI
client = OpenAI(base_url="https://api.belarel.com/v1", api_key=os.environ["BELAREL_API_KEY"])
response = client.responses.create( model="openai/gpt-5", reasoning={"effort": "low"}, input="Is 221 a prime number?",)print(response.output_text)print(response.usage.output_tokens_details.reasoning_tokens)import OpenAI from 'openai';
const client = new OpenAI({ baseURL: 'https://api.belarel.com/v1', apiKey: process.env.BELAREL_API_KEY });
const completion = await client.chat.completions.create({ model: 'openai/gpt-5', reasoning_effort: 'low', messages: [{ role: 'user', content: 'Is 221 a prime number?' }],});console.log(completion.choices[0].message.content);console.log(completion.usage?.completion_tokens_details?.reasoning_tokens);When the effort is not applied
Section titled “When the effort is not applied”Today the effort is applied to OpenAI models (openai/…). Other providers have no equivalent setting that Belarel translates.
Belarel does not drop the setting silently, and it does not refuse the call. The call runs at the model’s default, and belarel.policy.decisions says what happened:
"policy": { "decisions": [ { "rule": "org_api_catalog", "verdict": "allow", "detail": "residency any" }, { "rule": "reasoning_effort", "verdict": "warn", "detail": "not applied: vendor \"anthropic\" takes no effort setting" } ]}When the effort is applied, the same rule appears with "verdict": "allow" and the effort as detail, for example "detail": "low".
Read the reasoning
Section titled “Read the reasoning”Many models keep their reasoning private unless asked. Belarel asks OpenAI and Google models for a reasoning summary. Other models return reasoning only if their provider streams it by default. What you receive depends on the model.
| Dialect | Non-streamed | Streamed |
|---|---|---|
| Chat Completions | choices[0].message.reasoning_content (absent when there is none) |
delta.reasoning_content chunks |
| Responses | a reasoning output item, with its text in summary[].text (type: "summary_text") |
response.reasoning_summary_text.delta events |
You do not need to send reasoning back in the conversation. On the Responses API, reasoning items you include in input are accepted and skipped.
Reasoning tokens in usage
Section titled “Reasoning tokens in usage”Reasoning tokens are counted inside the output tokens, and the details field tells you how many there were:
| Dialect | Output tokens | Of which reasoning |
|---|---|---|
| Chat Completions | usage.completion_tokens |
usage.completion_tokens_details.reasoning_tokens |
| Responses | usage.output_tokens |
usage.output_tokens_details.reasoning_tokens |
A high effort can spend many tokens before the first word of the answer. Leave room in max_tokens (or max_output_tokens): the budget covers reasoning and answer together.
If the budget runs out, the answer is cut short and the response says so: finish_reason is length on Chat Completions, and on Responses status is incomplete with incomplete_details.reason: "max_output_tokens". A reasoning model can spend the whole budget thinking and return little or no text.
Long reasoning: stream it
Section titled “Long reasoning: stream it”Every call has a time limit of a little under 300 seconds. A non-streamed answer sends nothing until it is complete, so a long reasoning can reach that limit, or a proxy’s shorter idle timeout, before the first byte, and end in 504 request_timeout. With stream: true, reasoning and text reach you as they are generated, and Belarel sends a keep-alive comment every 15 seconds so the connection stays open. Use streaming for reasoning models with medium or high effort. See Streaming.

