Skip to content

Streaming

Set "stream": true on POST /v1/chat/completions or POST /v1/responses to receive the answer as server-sent events (Content-Type: text/event-stream) while it is generated. Without stream, or with false, you get one JSON object.

Both dialects end the stream with the usage and the belarel block, so you can read the cost of a streamed call without a second request.

Each event is a data: line holding a chat.completion.chunk. The stream goes in this order:

  1. A first chunk with delta.role: "assistant".
  2. Content chunks with delta.content. If the model returns visible reasoning, those parts arrive as delta.reasoning_content.
  3. A finish chunk with an empty delta and a finish_reason: stop, length (the output budget cut the answer), content_filter or tool_calls. See Chat Completions.
  4. A usage chunk with choices: [], usage and belarel.
  5. data: [DONE].
data: {"id":"chatcmpl-1","object":"chat.completion.chunk","created":1759300000,"model":"openai/gpt-5","choices":[{"index":0,"delta":{"role":"assistant"},"finish_reason":null}]}
data: {"id":"chatcmpl-1","object":"chat.completion.chunk","created":1759300000,"model":"openai/gpt-5","choices":[{"index":0,"delta":{"content":"Bonjour."},"finish_reason":null}]}
data: {"id":"chatcmpl-1","object":"chat.completion.chunk","created":1759300000,"model":"openai/gpt-5","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}
data: {"id":"chatcmpl-1","object":"chat.completion.chunk","created":1759300000,"model":"openai/gpt-5","choices":[],"usage":{"prompt_tokens":12,"completion_tokens":3,"total_tokens":15,"prompt_tokens_details":{"cached_tokens":0},"completion_tokens_details":{"reasoning_tokens":0}},"belarel":{"request_id":"req_0f1e","model_requested":"openai/gpt-5","model_served":"openai/gpt-5","cost":{"usd":0.00004,"credits":0.04,"source":"estimated","billing":"platform"},"policy":{"decisions":[{"rule":"org_api_catalog","verdict":"allow","detail":"residency any"}]},"data_mode":"standard"}}
data: [DONE]

The usage chunk is always sent. You do not need stream_options.include_usage: on Chat Completions it is accepted and has no effect. On Responses, stream_options is refused with 400 unsupported_parameter.

Other chunks you may see:

  • Tool calls. When the model calls one of your tools, each call arrives in a single chunk with delta.tool_calls, carrying the complete arguments string (not split into fragments). index counts the calls from 0. See Tool calling.
  • Citations. With web search, sources arrive as chunks with an empty delta and a top-level citations array of {url, title}.

Ignore fields you do not recognize. Belarel may add fields to chunks.

Each event has an event: line naming its type, and a data: line with the same type and a sequence_number that starts at 0 and increases by one per event.

event: response.created
data: {"type":"response.created","sequence_number":0,"response":{"id":"resp_0f1e","object":"response","status":"in_progress","model":"openai/gpt-5","output":[]}}
event: response.output_item.added
data: {"type":"response.output_item.added","sequence_number":2,"output_index":0,"item":{"type":"message","id":"msg_0f1e_0","status":"in_progress","role":"assistant","content":[]}}
event: response.output_text.delta
data: {"type":"response.output_text.delta","sequence_number":4,"item_id":"msg_0f1e_0","output_index":0,"content_index":0,"delta":"Bonjour."}
event: response.output_item.done
data: {"type":"response.output_item.done","sequence_number":7,"output_index":0,"item":{"type":"message","id":"msg_0f1e_0","status":"completed","role":"assistant","content":[{"type":"output_text","text":"Bonjour.","annotations":[]}]}}
event: response.completed
data: {"type":"response.completed","sequence_number":8,"response":{"id":"resp_0f1e","object":"response","status":"completed","model":"openai/gpt-5","output":[…],"usage":{"input_tokens":12,"input_tokens_details":{"cached_tokens":0},"output_tokens":3,"output_tokens_details":{"reasoning_tokens":0},"total_tokens":15},"belarel":{"request_id":"req_0f1e","cost":{"usd":0.00004,"credits":0.04,"source":"estimated","billing":"platform"},"data_mode":"standard"}}}

(Events 1, 3, 5 and 6 are response.in_progress, response.content_part.added, response.output_text.done and response.content_part.done.)

Event When
response.created, response.in_progress At the start.
response.output_item.added / .done An output item opens or closes: message, reasoning, function_call or web_search_call.
response.content_part.added / .done The text part of a message opens or closes.
response.output_text.delta / .done Message text.
response.reasoning_summary_part.added / .done, response.reasoning_summary_text.delta / .done Visible reasoning.
response.function_call_arguments.delta / .done A call to one of your tools. The arguments arrive whole, in one delta.
response.web_search_call.completed Web search ran for this answer.
response.completed The end. response holds the full output, usage and belarel.
response.incomplete The end, when the answer was cut short. response.status is incomplete and response.incomplete_details.reason is max_output_tokens or content_filter. It carries output, usage and belarel like response.completed.
response.failed The end, after a failure. See below.

Most refusals happen before the stream starts. They are ordinary JSON errors with a 4xx or 5xx status. See Errors.

Once the 200 headers are sent, a failure can no longer change the status code. It arrives in band, as the last event:

  • Chat Completions: an error event, then [DONE]. No finish chunk and no usage chunk are sent. The error has the same message, type and code the call would have returned as a JSON error.

    data: {"error":{"message":"The model could not be reached","type":"api_error","code":"service_unavailable"}}
    data: [DONE]
  • Responses: a response.failed event whose response.status is failed and whose response.error is {"code": "service_unavailable", "message": "…"}.

The code tells you what to do, as in Errors: service_unavailable is worth a retry, request_timeout means the call reached the time limit (see Long calls), and invalid_request means the provider refused the request itself.

The official OpenAI SDKs raise an exception when they read the Chat Completions error event. Treat a stream that ends without a finish chunk, or without response.completed or response.incomplete, as failed.

Every call has a time limit of a little under 300 seconds, streamed or not. When a call reaches it, Belarel stops the model and ends the call cleanly:

  • Not streamed: you get 504 request_timeout. The answer carries the header x-should-retry: false, so the official OpenAI SDKs do not replay a call that would only time out again.
  • Streamed: the stream ends with an in-band error whose code is request_timeout (Chat Completions), or a response.failed event with that code (Responses).

A non-streamed answer sends nothing until it is complete. A reasoning model can think for minutes before its first word, and a non-streamed call can then reach the limit with nothing to show. The limit for a non-streamed answer can also be lower than 300 seconds when a proxy with a shorter idle timeout sits in front of the API.

Stream long calls. With stream: true, tokens reach you as they are generated. Belarel also sends an SSE comment every 15 seconds for as long as the stream is open, so a silent stretch, such as a long reasoning, never looks like a dead connection:

: keep-alive

The official SDKs skip comment lines. If you parse the stream yourself, ignore lines that start with :.

A call that times out is settled at what the provider metered before it was stopped or, when nothing was metered, at the amount reserved for the call. See Cost and usage.

Terminal window
curl -N https://api.belarel.com/v1/chat/completions \
-H "Authorization: Bearer $BELAREL_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "openai/gpt-5", "stream": true,
"messages": [{"role": "user", "content": "Count to three."}]}'
# The last chunk before [DONE] carries "usage" and "belarel".

If the cost is not known yet when the stream closes, belarel.cost.source is pending and usd is null. Read the reconciled figure later with GET /v1/requests/{id} (scope usage:read). The request id is in the X-Request-Id response header, which you receive before the first event, and in belarel.request_id. See Request inspection.

If you close the connection early, generation stops and the call is still settled.