Skip to content

Responses API

POST /v1/responses speaks the OpenAI Responses dialect. It runs the same governed turn as Chat Completions, with the same admission, policies, cost and belarel block. Use it if your code already targets the Responses API, or if you want typed streaming events.

The key needs the inference scope.

Belarel does not store your responses. Each request must carry the whole conversation. Features that depend on stored state are refused with a 400, not silently ignored:

You send Result
store: true 400 unsupported_parameter, param: "store". store: false or no store is fine.
previous_response_id 400 unsupported_parameter, param: "previous_response_id"
An item_reference input item 400 unsupported_parameter
An input_file content part 400 unsupported_parameter. Send the text, or an image.
An input_image with file_id instead of image_url 400 invalid_request

Every response reports "store": false.

Terminal window
curl https://api.belarel.com/v1/responses \
-H "Authorization: Bearer $BELAREL_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "openai/gpt-5",
"instructions": "You answer in one sentence.",
"input": "What is a stateless API?",
"max_output_tokens": 200
}'
Field Notes
model Required.
input Required. A string, or a non-empty array of up to 500 items (see below).
instructions A string, sent as a system message ahead of input.
max_output_tokens The output budget, a positive integer. Capped by your key’s ceiling (400 max_tokens_exceeds_key_limit above it).
temperature A number from 0 to 2, forwarded to the model.
top_p A number from 0 to 1, forwarded to the model.
stream As in Chat Completions. See Streaming.
tools {"type": "function", "name", "description", "parameters"} tools that you execute, and {"type": "web_search"} (or web_search_preview). Any other tool type returns 400 unsupported_parameter.
tool_choice auto or none, forwarded to the model. With none, your tools stay declared and the model does not call them. Any other value returns 400 unsupported_parameter.
text.format Structured output. See Structured outputs.
reasoning.effort minimal, low, medium or high. See Reasoning.
metadata, belarel_tags Attribution tags (string values). See Tags and attribution.

user, parallel_tool_calls and service_tier are accepted with no effect, and so are store: false and truncation: "disabled". Fields nested inside an accepted object — such as text.verbosity or reasoning.summary — are not checked: Belarel ignores them without an error. These are refused by name:

You send Result
truncation: "auto" 400 unsupported_parameter. Send a conversation that fits the model’s context.
background: true 400 unsupported_parameter. Every response is computed while you wait.
top_logprobs 400 unsupported_parameter
stream_options, stop, seed, presence_penalty, frequency_penalty 400 unsupported_parameter. These work on Chat Completions, not here.
Any other top-level field not listed on this page, such as include, conversation or prompt 400 unsupported_parameter, param set to the field.

A forwarded parameter can still be dropped by the provider of the model. When that happens, belarel.ignored_parameters lists it — under its Chat Completions name: an ignored max_output_tokens appears as max_tokens, and text.format as response_format. Likewise, 400 model_lacks_structured_output names response_format in param, not text.format. See Chat Completions.

Item Notes
message role is user, assistant, system or developer (developer is treated as system). content is a string or an array of input_text, output_text and input_image parts. An item with a role and no type is read as a message.
function_call A call the model made earlier. Needs call_id and name.
function_call_output Your tool’s result. Needs call_id. A non-string output is sent as JSON.
reasoning, web_search_call Accepted and skipped, so you can send back the output array you received.

Because nothing is stored, you keep the history and send it back on every turn. The simplest way is to append the output items you received, then your next message:

history = [{"role": "user", "content": "Name a prime number above 100."}]
first = client.responses.create(model="openai/gpt-5", input=history)
history += first.output # message, reasoning and function_call items
history.append({"role": "user", "content": "And the next one?"})
second = client.responses.create(model="openai/gpt-5", input=history)
print(second.output_text)

The same pattern works for tools: append the function_call item, then a function_call_output item with the same call_id. See Tool calling.

{
"id": "resp_0f1e…",
"object": "response",
"created_at": 1759300000,
"status": "completed",
"model": "openai/gpt-5",
"output": [
{
"type": "message",
"id": "msg_0f1e…_0",
"status": "completed",
"role": "assistant",
"content": [{ "type": "output_text", "text": "An API that keeps no session between calls.", "annotations": [] }]
}
],
"error": null,
"incomplete_details": null,
"parallel_tool_calls": true,
"store": false,
"usage": {
"input_tokens": 24,
"input_tokens_details": { "cached_tokens": 0 },
"output_tokens": 11,
"output_tokens_details": { "reasoning_tokens": 0 },
"total_tokens": 35
},
"belarel": { "request_id": "req_0f1e…", "cost": { "usd": 0.00004, "credits": 0.04, "source": "estimated", "billing": "platform" }, "…": "…" }
}
  • output can hold message items (with url_citation annotations when web search ran), reasoning items (summary_text), function_call items and a web_search_call item.
  • The response id and the request id name the same call: resp_<hex> matches req_<hex> in X-Request-Id and belarel.request_id. GET /v1/requests/{id} accepts either form.
  • belarel is the same block as in Chat Completions.
  • status is completed, or incomplete when the answer was cut short. Then incomplete_details.reason is max_output_tokens (the output budget ran out) or content_filter (the provider filtered the answer), and the cut message item also has "status": "incomplete".