Cost & usage
Every successful answer tells you what it cost, in the answer itself: an OpenAI-shaped usage object and a belarel block with the request id, the model that served it, the cost in USD and credits, the policy decisions, and any parameter the model ignored. For totals over time, GET /v1/usage aggregates the same figures.
The belarel block
Section titled “The belarel block”"belarel": { "request_id": "req_6f1c0a7e2b9d4c3f8e5a1b2c3d4e5f60", "model_requested": "<model-id>", "model_served": "<model-id>", "cost": { "usd": 0.000412, "credits": 0.412, "source": "estimated", "billing": "platform" }, "policy": { "decisions": [{ "rule": "org_api_catalog", "verdict": "allow", "detail": "residency any" }] }, "data_mode": "standard", "ignored_parameters": ["seed"]}| Field | Meaning |
|---|---|
request_id |
The call’s id — the same value as the X-Request-Id header. |
model_requested |
The model you sent. |
model_served |
The model that served the call. |
cost.usd |
The call’s cost in USD at your organization’s sale price. null while pending. |
cost.credits |
The same cost in credits, the unit of your organization’s balance. null while pending. |
cost.source |
estimated: computed when the call closed, from the usage reported for it. pending: not known within the wait, see below. metered is also defined; treat it like estimated. |
cost.billing |
platform: billed by Belarel. byok: served on your organization’s own provider key — usd and credits are Belarel’s fee only. sandbox: never billed, always 0. |
policy.decisions |
The rules that let the call through. See Policies & catalog. |
data_mode |
standard or zero_retention. See Data modes. |
ignored_parameters |
Present only when the model’s provider dropped parameters you sent and Belarel forwarded — for example seed, the penalties, or top_p next to temperature. Listed by their OpenAI names. |
When cost.source is pending
Section titled “When cost.source is pending”The answer waits up to 2 seconds after the model finishes for the cost. If it is not ready — or no cost could be computed — usd and credits are null and source is pending. The call is still settled; read its cost later with GET /v1/requests/{id}, where cost.source moves to estimated, then to reconciled once the provider’s real cost is known. If no cost could be computed at all, cost.source stays pending there for good; the call then counts against your caps at its estimate, which you can read in cost.settled_usd. See Request inspection.
Usage, including cached and reasoning tokens
Section titled “Usage, including cached and reasoning tokens”Chat Completions answers carry the OpenAI usage shape, completed with cached input tokens and reasoning tokens:
"usage": { "prompt_tokens": 1250, "completion_tokens": 310, "total_tokens": 1560, "prompt_tokens_details": { "cached_tokens": 1024 }, "completion_tokens_details": { "reasoning_tokens": 128 }}Responses answers use the Responses names: input_tokens, input_tokens_details.cached_tokens, output_tokens, output_tokens_details.reasoning_tokens, total_tokens. Embeddings answers report prompt_tokens and total_tokens. A detail is 0 when the model reported none.
Streaming
Section titled “Streaming”You do not need stream_options.include_usage: a streamed answer always ends with its usage and its cost. (Chat Completions accepts the field with no effect; Responses refuses it.) While the model works, the stream may carry : keep-alive comment lines every 15 seconds; SDKs ignore them, and so should a hand-written SSE parser.
- Chat Completions — after the chunk that carries
finish_reason, and beforedata: [DONE], Belarel sends one more chunk withchoices: [],usageandbelarel. - Responses — the
responseobject of the last event carriesusageandbelarel:response.completed, orresponse.incompletewhen the answer was cut short (for example bymax_output_tokens).
curl -N https://api.belarel.com/v1/chat/completions \ -H "Authorization: Bearer $BELAREL_API_KEY" \ -H "Content-Type: application/json" \ -d '{"model": "<model-id>", "stream": true, "messages": [{"role": "user", "content": "Hello"}]}'# ... the last data line before [DONE] holds "choices":[], "usage" and "belarel"import osfrom openai import OpenAI
client = OpenAI(base_url="https://api.belarel.com/v1", api_key=os.environ["BELAREL_API_KEY"])stream = client.chat.completions.create( model="<model-id>", stream=True, messages=[{"role": "user", "content": "Hello"}],)for chunk in stream: if chunk.choices: print(chunk.choices[0].delta.content or "", end="") else: # the closing chunk belarel = (chunk.model_extra or {}).get("belarel") print("\n", chunk.usage, belarel["cost"] if belarel else None)import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.belarel.com/v1", apiKey: process.env.BELAREL_API_KEY });const stream = await client.chat.completions.create({ model: "<model-id>", stream: true, messages: [{ role: "user", content: "Hello" }],});for await (const chunk of stream) { if (chunk.choices.length > 0) { process.stdout.write(chunk.choices[0].delta.content ?? ""); } else { const belarel = (chunk as any).belarel; // the closing chunk console.log("\n", chunk.usage, belarel?.cost); }}Aggregated usage: GET /v1/usage
Section titled “Aggregated usage: GET /v1/usage”GET /v1/usage (scope usage:read) aggregates the calls of the key and of its sub-keys. It never returns content.
| Parameter | Default | Values |
|---|---|---|
from |
30 days before to |
ISO 8601 date or Unix seconds |
to |
now | ISO 8601 date or Unix seconds; the window holds at most 93 days |
group_by |
total |
total, key, model, day (UTC), or tag:<name> |
curl "https://api.belarel.com/v1/usage?from=2026-09-01&to=2026-10-01&group_by=tag:customer" \ -H "Authorization: Bearer $BELAREL_API_KEY"{ "object": "list", "from": "2026-09-01T00:00:00.000Z", "to": "2026-10-01T00:00:00.000Z", "group_by": "tag:customer", "data": [ { "group": "acme", "label": null, "requests": 1840, "failed_requests": 3, "input_tokens": 2210400, "output_tokens": 401230, "cached_tokens": 880000, "reasoning_tokens": 0, "credits": 1520.4, "cost_usd": 1.5204, "reconciled_requests": 1790 } ], "belarel": { "request_id": "req_…" }}group is the key id, model, day (YYYY-MM-DD), tag value or total; label is the key’s name when you group by key. Calls without the tag you group by fall into a group whose value is null. reconciled_requests counts the calls whose provider cost has been confirmed — the totals settle as that number catches up with requests. Calls refused at admission — key, policy, caps, rate limits, credits — are not counted. An invalid window or group_by answers 400 invalid_request.
Credits and USD
Section titled “Credits and USD”USD is how your spend caps are expressed and how cost.usd reports a call, at your organization’s sale price. Credits are the unit of your organization’s balance: each call’s cost.credits is deducted from it, and GET /v1/credits shows what remains. The conversion between the two comes from your organization’s price list. See Budgets, caps & credits.

