Skip to content

Cost & usage

Every successful answer tells you what it cost, in the answer itself: an OpenAI-shaped usage object and a belarel block with the request id, the model that served it, the cost in USD and credits, the policy decisions, and any parameter the model ignored. For totals over time, GET /v1/usage aggregates the same figures.

"belarel": {
"request_id": "req_6f1c0a7e2b9d4c3f8e5a1b2c3d4e5f60",
"model_requested": "<model-id>",
"model_served": "<model-id>",
"cost": { "usd": 0.000412, "credits": 0.412, "source": "estimated", "billing": "platform" },
"policy": { "decisions": [{ "rule": "org_api_catalog", "verdict": "allow", "detail": "residency any" }] },
"data_mode": "standard",
"ignored_parameters": ["seed"]
}
Field Meaning
request_id The call’s id — the same value as the X-Request-Id header.
model_requested The model you sent.
model_served The model that served the call.
cost.usd The call’s cost in USD at your organization’s sale price. null while pending.
cost.credits The same cost in credits, the unit of your organization’s balance. null while pending.
cost.source estimated: computed when the call closed, from the usage reported for it. pending: not known within the wait, see below. metered is also defined; treat it like estimated.
cost.billing platform: billed by Belarel. byok: served on your organization’s own provider key — usd and credits are Belarel’s fee only. sandbox: never billed, always 0.
policy.decisions The rules that let the call through. See Policies & catalog.
data_mode standard or zero_retention. See Data modes.
ignored_parameters Present only when the model’s provider dropped parameters you sent and Belarel forwarded — for example seed, the penalties, or top_p next to temperature. Listed by their OpenAI names.

The answer waits up to 2 seconds after the model finishes for the cost. If it is not ready — or no cost could be computed — usd and credits are null and source is pending. The call is still settled; read its cost later with GET /v1/requests/{id}, where cost.source moves to estimated, then to reconciled once the provider’s real cost is known. If no cost could be computed at all, cost.source stays pending there for good; the call then counts against your caps at its estimate, which you can read in cost.settled_usd. See Request inspection.

Usage, including cached and reasoning tokens

Section titled “Usage, including cached and reasoning tokens”

Chat Completions answers carry the OpenAI usage shape, completed with cached input tokens and reasoning tokens:

"usage": {
"prompt_tokens": 1250,
"completion_tokens": 310,
"total_tokens": 1560,
"prompt_tokens_details": { "cached_tokens": 1024 },
"completion_tokens_details": { "reasoning_tokens": 128 }
}

Responses answers use the Responses names: input_tokens, input_tokens_details.cached_tokens, output_tokens, output_tokens_details.reasoning_tokens, total_tokens. Embeddings answers report prompt_tokens and total_tokens. A detail is 0 when the model reported none.

You do not need stream_options.include_usage: a streamed answer always ends with its usage and its cost. (Chat Completions accepts the field with no effect; Responses refuses it.) While the model works, the stream may carry : keep-alive comment lines every 15 seconds; SDKs ignore them, and so should a hand-written SSE parser.

  • Chat Completions — after the chunk that carries finish_reason, and before data: [DONE], Belarel sends one more chunk with choices: [], usage and belarel.
  • Responses — the response object of the last event carries usage and belarel: response.completed, or response.incomplete when the answer was cut short (for example by max_output_tokens).
Terminal window
curl -N https://api.belarel.com/v1/chat/completions \
-H "Authorization: Bearer $BELAREL_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "<model-id>", "stream": true, "messages": [{"role": "user", "content": "Hello"}]}'
# ... the last data line before [DONE] holds "choices":[], "usage" and "belarel"

GET /v1/usage (scope usage:read) aggregates the calls of the key and of its sub-keys. It never returns content.

Parameter Default Values
from 30 days before to ISO 8601 date or Unix seconds
to now ISO 8601 date or Unix seconds; the window holds at most 93 days
group_by total total, key, model, day (UTC), or tag:<name>
Terminal window
curl "https://api.belarel.com/v1/usage?from=2026-09-01&to=2026-10-01&group_by=tag:customer" \
-H "Authorization: Bearer $BELAREL_API_KEY"
{
"object": "list",
"from": "2026-09-01T00:00:00.000Z",
"to": "2026-10-01T00:00:00.000Z",
"group_by": "tag:customer",
"data": [
{ "group": "acme", "label": null, "requests": 1840, "failed_requests": 3,
"input_tokens": 2210400, "output_tokens": 401230, "cached_tokens": 880000, "reasoning_tokens": 0,
"credits": 1520.4, "cost_usd": 1.5204, "reconciled_requests": 1790 }
],
"belarel": { "request_id": "req_…" }
}

group is the key id, model, day (YYYY-MM-DD), tag value or total; label is the key’s name when you group by key. Calls without the tag you group by fall into a group whose value is null. reconciled_requests counts the calls whose provider cost has been confirmed — the totals settle as that number catches up with requests. Calls refused at admission — key, policy, caps, rate limits, credits — are not counted. An invalid window or group_by answers 400 invalid_request.

USD is how your spend caps are expressed and how cost.usd reports a call, at your organization’s sale price. Credits are the unit of your organization’s balance: each call’s cost.credits is deducted from it, and GET /v1/credits shows what remains. The conversion between the two comes from your organization’s price list. See Budgets, caps & credits.