Streaming
Set "stream": true on POST /v1/chat/completions or POST /v1/responses to receive the answer as server-sent events (Content-Type: text/event-stream) while it is generated. Without stream, or with false, you get one JSON object.
Both dialects end the stream with the usage and the belarel block, so you can read the cost of a streamed call without a second request.
Chat Completions
Section titled “Chat Completions”Each event is a data: line holding a chat.completion.chunk. The stream goes in this order:
- A first chunk with
delta.role: "assistant". - Content chunks with
delta.content. If the model returns visible reasoning, those parts arrive asdelta.reasoning_content. - A finish chunk with an empty
deltaand afinish_reason:stop,length(the output budget cut the answer),content_filterortool_calls. See Chat Completions. - A usage chunk with
choices: [],usageandbelarel. data: [DONE].
data: {"id":"chatcmpl-1","object":"chat.completion.chunk","created":1759300000,"model":"openai/gpt-5","choices":[{"index":0,"delta":{"role":"assistant"},"finish_reason":null}]}
data: {"id":"chatcmpl-1","object":"chat.completion.chunk","created":1759300000,"model":"openai/gpt-5","choices":[{"index":0,"delta":{"content":"Bonjour."},"finish_reason":null}]}
data: {"id":"chatcmpl-1","object":"chat.completion.chunk","created":1759300000,"model":"openai/gpt-5","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}
data: {"id":"chatcmpl-1","object":"chat.completion.chunk","created":1759300000,"model":"openai/gpt-5","choices":[],"usage":{"prompt_tokens":12,"completion_tokens":3,"total_tokens":15,"prompt_tokens_details":{"cached_tokens":0},"completion_tokens_details":{"reasoning_tokens":0}},"belarel":{"request_id":"req_0f1e","model_requested":"openai/gpt-5","model_served":"openai/gpt-5","cost":{"usd":0.00004,"credits":0.04,"source":"estimated","billing":"platform"},"policy":{"decisions":[{"rule":"org_api_catalog","verdict":"allow","detail":"residency any"}]},"data_mode":"standard"}}
data: [DONE]The usage chunk is always sent. You do not need stream_options.include_usage: on Chat Completions it is accepted and has no effect. On Responses, stream_options is refused with 400 unsupported_parameter.
Other chunks you may see:
- Tool calls. When the model calls one of your tools, each call arrives in a single chunk with
delta.tool_calls, carrying the completeargumentsstring (not split into fragments).indexcounts the calls from 0. See Tool calling. - Citations. With web search, sources arrive as chunks with an empty
deltaand a top-levelcitationsarray of{url, title}.
Ignore fields you do not recognize. Belarel may add fields to chunks.
Responses API
Section titled “Responses API”Each event has an event: line naming its type, and a data: line with the same type and a sequence_number that starts at 0 and increases by one per event.
event: response.createddata: {"type":"response.created","sequence_number":0,"response":{"id":"resp_0f1e","object":"response","status":"in_progress","model":"openai/gpt-5","output":[]}}
event: response.output_item.addeddata: {"type":"response.output_item.added","sequence_number":2,"output_index":0,"item":{"type":"message","id":"msg_0f1e_0","status":"in_progress","role":"assistant","content":[]}}
event: response.output_text.deltadata: {"type":"response.output_text.delta","sequence_number":4,"item_id":"msg_0f1e_0","output_index":0,"content_index":0,"delta":"Bonjour."}
event: response.output_item.donedata: {"type":"response.output_item.done","sequence_number":7,"output_index":0,"item":{"type":"message","id":"msg_0f1e_0","status":"completed","role":"assistant","content":[{"type":"output_text","text":"Bonjour.","annotations":[]}]}}
event: response.completeddata: {"type":"response.completed","sequence_number":8,"response":{"id":"resp_0f1e","object":"response","status":"completed","model":"openai/gpt-5","output":[…],"usage":{"input_tokens":12,"input_tokens_details":{"cached_tokens":0},"output_tokens":3,"output_tokens_details":{"reasoning_tokens":0},"total_tokens":15},"belarel":{"request_id":"req_0f1e","cost":{"usd":0.00004,"credits":0.04,"source":"estimated","billing":"platform"},"data_mode":"standard"}}}(Events 1, 3, 5 and 6 are response.in_progress, response.content_part.added, response.output_text.done and response.content_part.done.)
| Event | When |
|---|---|
response.created, response.in_progress |
At the start. |
response.output_item.added / .done |
An output item opens or closes: message, reasoning, function_call or web_search_call. |
response.content_part.added / .done |
The text part of a message opens or closes. |
response.output_text.delta / .done |
Message text. |
response.reasoning_summary_part.added / .done, response.reasoning_summary_text.delta / .done |
Visible reasoning. |
response.function_call_arguments.delta / .done |
A call to one of your tools. The arguments arrive whole, in one delta. |
response.web_search_call.completed |
Web search ran for this answer. |
response.completed |
The end. response holds the full output, usage and belarel. |
response.incomplete |
The end, when the answer was cut short. response.status is incomplete and response.incomplete_details.reason is max_output_tokens or content_filter. It carries output, usage and belarel like response.completed. |
response.failed |
The end, after a failure. See below. |
Errors after the stream has started
Section titled “Errors after the stream has started”Most refusals happen before the stream starts. They are ordinary JSON errors with a 4xx or 5xx status. See Errors.
Once the 200 headers are sent, a failure can no longer change the status code. It arrives in band, as the last event:
-
Chat Completions: an
errorevent, then[DONE]. No finish chunk and no usage chunk are sent. The error has the samemessage,typeandcodethe call would have returned as a JSON error.data: {"error":{"message":"The model could not be reached","type":"api_error","code":"service_unavailable"}}data: [DONE] -
Responses: a
response.failedevent whoseresponse.statusisfailedand whoseresponse.erroris{"code": "service_unavailable", "message": "…"}.
The code tells you what to do, as in Errors: service_unavailable is worth a retry, request_timeout means the call reached the time limit (see Long calls), and invalid_request means the provider refused the request itself.
The official OpenAI SDKs raise an exception when they read the Chat Completions error event. Treat a stream that ends without a finish chunk, or without response.completed or response.incomplete, as failed.
Long calls
Section titled “Long calls”Every call has a time limit of a little under 300 seconds, streamed or not. When a call reaches it, Belarel stops the model and ends the call cleanly:
- Not streamed: you get
504 request_timeout. The answer carries the headerx-should-retry: false, so the official OpenAI SDKs do not replay a call that would only time out again. - Streamed: the stream ends with an in-band error whose
codeisrequest_timeout(Chat Completions), or aresponse.failedevent with that code (Responses).
A non-streamed answer sends nothing until it is complete. A reasoning model can think for minutes before its first word, and a non-streamed call can then reach the limit with nothing to show. The limit for a non-streamed answer can also be lower than 300 seconds when a proxy with a shorter idle timeout sits in front of the API.
Stream long calls. With stream: true, tokens reach you as they are generated. Belarel also sends an SSE comment every 15 seconds for as long as the stream is open, so a silent stretch, such as a long reasoning, never looks like a dead connection:
: keep-aliveThe official SDKs skip comment lines. If you parse the stream yourself, ignore lines that start with :.
A call that times out is settled at what the provider metered before it was stopped or, when nothing was metered, at the amount reserved for the call. See Cost and usage.
Read the cost of a streamed call
Section titled “Read the cost of a streamed call”curl -N https://api.belarel.com/v1/chat/completions \ -H "Authorization: Bearer $BELAREL_API_KEY" \ -H "Content-Type: application/json" \ -d '{"model": "openai/gpt-5", "stream": true, "messages": [{"role": "user", "content": "Count to three."}]}'# The last chunk before [DONE] carries "usage" and "belarel".import osfrom openai import OpenAI
client = OpenAI(base_url="https://api.belarel.com/v1", api_key=os.environ["BELAREL_API_KEY"])
stream = client.chat.completions.create( model="openai/gpt-5", messages=[{"role": "user", "content": "Count to three."}], stream=True,)belarel = Nonefor chunk in stream: if chunk.choices and chunk.choices[0].delta.content: print(chunk.choices[0].delta.content, end="") if chunk.usage: belarel = getattr(chunk, "belarel", None)print("\n", belarel["cost"] if belarel else None)import OpenAI from 'openai';
const client = new OpenAI({ baseURL: 'https://api.belarel.com/v1', apiKey: process.env.BELAREL_API_KEY });
const stream = await client.chat.completions.create({ model: 'openai/gpt-5', messages: [{ role: 'user', content: 'Count to three.' }], stream: true,});let belarel: { cost: unknown } | undefined;for await (const chunk of stream) { process.stdout.write(chunk.choices[0]?.delta?.content ?? ''); if (chunk.usage) belarel = (chunk as unknown as { belarel: { cost: unknown } }).belarel;}console.log('\n', belarel?.cost);If the cost is not known yet when the stream closes, belarel.cost.source is pending and usd is null. Read the reconciled figure later with GET /v1/requests/{id} (scope usage:read). The request id is in the X-Request-Id response header, which you receive before the first event, and in belarel.request_id. See Request inspection.
If you close the connection early, generation stops and the call is still settled.

