Skip to content

Data modes

Every API key carries a data mode. It decides how the content of the key’s calls — your prompts and the model’s answers — is handled, both by Belarel and by the model providers your calls are routed to. You pick it when you create the key; every call made with that key follows it, without any per-request flag.

Mode Status What it means
standard Available The default. Calls go to any model your organization enabled. The request journal may keep the content of the call for support and debugging, subject to your organization’s retention period.
zero_retention Available Belarel does not keep the content of the call, and the call is routed only to providers that offer zero data retention and do not train on prompts. Models that cannot be served this way are refused. See Zero data retention.
confidential Coming soon Reserved for attested confidential execution. It is part of the API contract, but cannot be selected yet, and calls are refused with 501 not_implemented. See Confidential inference.

What changes between standard and zero_retention

Section titled “What changes between standard and zero_retention”
standard zero_retention
Prompt and output in the request journal May be stored Never stored — no text, no hash, no image URL
Keyed fingerprint of the request body, for Idempotency-Key Kept up to 48 hours when the header is sent Same — omit the header to avoid it
Routing Any provider serving the model Only providers with zero retention and no training on prompts
Models you can call Every model your organization enabled for the key Only models that can be served with zero retention; others answer 400 model_not_zdr
Idempotency-Key replay of a JSON answer Replayed for 24 hours, then deleted within the next day — stored up to 48 hours Never stored; a retry answers 409 idempotency_conflict
Tokens, cost, model, status, tags, timing Recorded Recorded
Billing Same Same

Billing, spend caps, rate limits and policies work identically in both modes: they rely on token counts and costs, never on content.

  • Use standard when you want the broadest model choice and you are comfortable with call content being kept for the retention period your organization sets.
  • Use zero_retention when the content of your calls must not persist anywhere you do not control — personal data, regulated content, customer documents. Check the model list first: in GET /v1/models, belarel.zero_retention_available: true marks models that can be served this way.
  • Do not plan on confidential yet. It is listed here so that your code can recognize the value.

You choose the data mode when you create a key in your organization dashboard, under Governance → API keys (the Zero data retention option). A key’s data mode is fixed at creation: there is no way to change it afterward. To switch modes, create a new key and revoke the old one.

Sub-keys created through POST /v1/keys follow one rule: a sub-key can be stricter than its parent, never looser.

  • A sub-key of a zero_retention key is always zero_retention, whatever you ask for.
  • A sub-key of a standard key is standard unless you pass "data_mode": "zero_retention".

The public representation of a key (GET /v1/key, GET /v1/keys) includes its data_mode.

Every successful response carries the mode the call was served under, in the belarel block:

{
"id": "chatcmpl-…",
"object": "chat.completion",
"choices": [ … ],
"usage": { … },
"belarel": {
"request_id": "req_6f1c…",
"model_requested": "<model-id>",
"model_served": "<model-id>",
"cost": { "usd": 0.00042, "credits": 0.42, "source": "estimated", "billing": "platform" },
"policy": { "decisions": [ … ] },
"data_mode": "zero_retention"
}
}

For a streamed Chat Completions answer, the same belarel block arrives in the final usage chunk, just before data: [DONE]; for a streamed Responses answer, it is on the response carried by the last event, response.completed or response.incomplete. If you need proof, per call, that a request ran under zero_retention, log belarel.data_mode together with belarel.request_id on your side.