Responses API
POST /v1/responses speaks the OpenAI Responses dialect. It runs the same governed turn as Chat Completions, with the same admission, policies, cost and belarel block. Use it if your code already targets the Responses API, or if you want typed streaming events.
The key needs the inference scope.
Stateless by design
Section titled “Stateless by design”Belarel does not store your responses. Each request must carry the whole conversation. Features that depend on stored state are refused with a 400, not silently ignored:
| You send | Result |
|---|---|
store: true |
400 unsupported_parameter, param: "store". store: false or no store is fine. |
previous_response_id |
400 unsupported_parameter, param: "previous_response_id" |
An item_reference input item |
400 unsupported_parameter |
An input_file content part |
400 unsupported_parameter. Send the text, or an image. |
An input_image with file_id instead of image_url |
400 invalid_request |
Every response reports "store": false.
A first call
Section titled “A first call”curl https://api.belarel.com/v1/responses \ -H "Authorization: Bearer $BELAREL_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "openai/gpt-5", "instructions": "You answer in one sentence.", "input": "What is a stateless API?", "max_output_tokens": 200 }'import osfrom openai import OpenAI
client = OpenAI(base_url="https://api.belarel.com/v1", api_key=os.environ["BELAREL_API_KEY"])
response = client.responses.create( model="openai/gpt-5", instructions="You answer in one sentence.", input="What is a stateless API?", max_output_tokens=200,)print(response.output_text)import OpenAI from 'openai';
const client = new OpenAI({ baseURL: 'https://api.belarel.com/v1', apiKey: process.env.BELAREL_API_KEY });
const response = await client.responses.create({ model: 'openai/gpt-5', instructions: 'You answer in one sentence.', input: 'What is a stateless API?', max_output_tokens: 200,});console.log(response.output_text);What the request supports
Section titled “What the request supports”| Field | Notes |
|---|---|
model |
Required. |
input |
Required. A string, or a non-empty array of up to 500 items (see below). |
instructions |
A string, sent as a system message ahead of input. |
max_output_tokens |
The output budget, a positive integer. Capped by your key’s ceiling (400 max_tokens_exceeds_key_limit above it). |
temperature |
A number from 0 to 2, forwarded to the model. |
top_p |
A number from 0 to 1, forwarded to the model. |
stream |
As in Chat Completions. See Streaming. |
tools |
{"type": "function", "name", "description", "parameters"} tools that you execute, and {"type": "web_search"} (or web_search_preview). Any other tool type returns 400 unsupported_parameter. |
tool_choice |
auto or none, forwarded to the model. With none, your tools stay declared and the model does not call them. Any other value returns 400 unsupported_parameter. |
text.format |
Structured output. See Structured outputs. |
reasoning.effort |
minimal, low, medium or high. See Reasoning. |
metadata, belarel_tags |
Attribution tags (string values). See Tags and attribution. |
user, parallel_tool_calls and service_tier are accepted with no effect, and so are store: false and truncation: "disabled". Fields nested inside an accepted object — such as text.verbosity or reasoning.summary — are not checked: Belarel ignores them without an error. These are refused by name:
| You send | Result |
|---|---|
truncation: "auto" |
400 unsupported_parameter. Send a conversation that fits the model’s context. |
background: true |
400 unsupported_parameter. Every response is computed while you wait. |
top_logprobs |
400 unsupported_parameter |
stream_options, stop, seed, presence_penalty, frequency_penalty |
400 unsupported_parameter. These work on Chat Completions, not here. |
Any other top-level field not listed on this page, such as include, conversation or prompt |
400 unsupported_parameter, param set to the field. |
A forwarded parameter can still be dropped by the provider of the model. When that happens, belarel.ignored_parameters lists it — under its Chat Completions name: an ignored max_output_tokens appears as max_tokens, and text.format as response_format. Likewise, 400 model_lacks_structured_output names response_format in param, not text.format. See Chat Completions.
Input items
Section titled “Input items”| Item | Notes |
|---|---|
message |
role is user, assistant, system or developer (developer is treated as system). content is a string or an array of input_text, output_text and input_image parts. An item with a role and no type is read as a message. |
function_call |
A call the model made earlier. Needs call_id and name. |
function_call_output |
Your tool’s result. Needs call_id. A non-string output is sent as JSON. |
reasoning, web_search_call |
Accepted and skipped, so you can send back the output array you received. |
Carry the conversation yourself
Section titled “Carry the conversation yourself”Because nothing is stored, you keep the history and send it back on every turn. The simplest way is to append the output items you received, then your next message:
history = [{"role": "user", "content": "Name a prime number above 100."}]first = client.responses.create(model="openai/gpt-5", input=history)
history += first.output # message, reasoning and function_call itemshistory.append({"role": "user", "content": "And the next one?"})second = client.responses.create(model="openai/gpt-5", input=history)print(second.output_text)The same pattern works for tools: append the function_call item, then a function_call_output item with the same call_id. See Tool calling.
The response object
Section titled “The response object”{ "id": "resp_0f1e…", "object": "response", "created_at": 1759300000, "status": "completed", "model": "openai/gpt-5", "output": [ { "type": "message", "id": "msg_0f1e…_0", "status": "completed", "role": "assistant", "content": [{ "type": "output_text", "text": "An API that keeps no session between calls.", "annotations": [] }] } ], "error": null, "incomplete_details": null, "parallel_tool_calls": true, "store": false, "usage": { "input_tokens": 24, "input_tokens_details": { "cached_tokens": 0 }, "output_tokens": 11, "output_tokens_details": { "reasoning_tokens": 0 }, "total_tokens": 35 }, "belarel": { "request_id": "req_0f1e…", "cost": { "usd": 0.00004, "credits": 0.04, "source": "estimated", "billing": "platform" }, "…": "…" }}outputcan holdmessageitems (withurl_citationannotations when web search ran),reasoningitems (summary_text),function_callitems and aweb_search_callitem.- The response id and the request id name the same call:
resp_<hex>matchesreq_<hex>inX-Request-Idandbelarel.request_id.GET /v1/requests/{id}accepts either form. belarelis the same block as in Chat Completions.statusiscompleted, orincompletewhen the answer was cut short. Thenincomplete_details.reasonismax_output_tokens(the output budget ran out) orcontent_filter(the provider filtered the answer), and the cutmessageitem also has"status": "incomplete".

