Zero data retention
A key in zero_retention mode is for workloads whose content must not persist: Belarel does not keep the prompt or the answer, and the call is only routed to model providers that offer zero data retention and do not train on prompts. This page describes exactly what that covers, and what it does not.
What zero retention does
Section titled “What zero retention does”For every call made with a zero_retention key, on POST /v1/chat/completions, POST /v1/responses and POST /v1/embeddings:
- Routing is restricted. The call carries two routing requirements: zero data retention, and no training on prompts. Only providers that satisfy both may serve it.
- A call that cannot be honoured is refused, not downgraded. If the model is not served through a route that supports zero retention, or no provider of that model can satisfy the requirements, the call fails with
400 model_not_zdr. It is never silently served by a provider that ignores the requirement. - No content is journaled. The request journal keeps no prompt, no output, no hash of either, no image URL and no content-derived cache key — regardless of how journaling is configured for
standardkeys. - No content is kept for retries. A JSON answer is never stored for
Idempotency-Keyreplay. Retrying with the same key returns409 idempotency_conflictnaming the original request id, instead of the answer.
The answer itself is returned to you as usual; Belarel does not keep a copy.
Which models qualify
Section titled “Which models qualify”GET /v1/models marks each model with belarel.zero_retention_available:
{ "id": "<model-id>", "object": "model", "belarel": { "category": "language", "zero_retention_available": true, "byok_available": false }}true means the model is served through a route that supports zero retention. Treat it as a prerequisite, not a guarantee: at call time, if no provider of that model can meet the requirements, the call is still refused with model_not_zdr. The list returned to a zero_retention key is not filtered on this flag, so filter it yourself.
Sandbox keys (bel_test_…) never reach a provider, so a sandbox zero_retention key accepts any model in its catalog. Test the real model list with a live key.
What is still recorded
Section titled “What is still recorded”Zero retention is about content. To bill you, enforce your caps and let you inspect your usage, each call still produces a record without content:
| Recorded | Not recorded |
|---|---|
| Request id, key id, timestamp, duration | Messages, system prompt, input text |
| Model requested and served, and the provider it was routed to | Model output |
| Token counts (input, output, cached, reasoning) | Hashes of the prompt or output in the request journal |
| Cost (USD and credits), billing ledger entry | Image URLs |
| HTTP status, success or failure, error message | Content-derived cache keys |
Tags (including the metadata field, and the automatic app and app_url tags from X-Title and HTTP-Referer) |
Stored response for idempotent replay |
Keyed fingerprint (HMAC-SHA256) of the request body, only when you send Idempotency-Key (up to 48 hours) |
GET /v1/requests/{id} returns the same record for a zero_retention call as for any other: status, model, usage, cost and tags — never content.
Using a zero retention key
Section titled “Using a zero retention key”Nothing changes in your code. The mode belongs to the key, so you only swap the key.
curl https://api.belarel.com/v1/chat/completions \ -H "Authorization: Bearer $BELAREL_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "<model-id>", "messages": [{ "role": "user", "content": "Summarize this clinical note: …" }] }'import osfrom openai import OpenAI
client = OpenAI(base_url="https://api.belarel.com/v1", api_key=os.environ["BELAREL_API_KEY"])
completion = client.chat.completions.create( model="<model-id>", messages=[{"role": "user", "content": "Summarize this clinical note: …"}],)print(completion.model_extra["belarel"]["data_mode"]) # "zero_retention"import OpenAI from 'openai';
const client = new OpenAI({ baseURL: 'https://api.belarel.com/v1', apiKey: process.env.BELAREL_API_KEY });
const completion = await client.chat.completions.create({ model: '<model-id>', messages: [{ role: 'user', content: 'Summarize this clinical note: …' }],});console.log((completion as any).belarel.data_mode); // "zero_retention"Errors specific to zero retention
Section titled “Errors specific to zero retention”| Code | HTTP | Meaning | What to do |
|---|---|---|---|
model_not_zdr |
400 | The model cannot be served with zero retention. | Pick a model with zero_retention_available: true. Do not retry the same model. |
idempotency_conflict |
409 | You retried with an Idempotency-Key already used; the answer was not stored. |
Inspect the original call with GET /v1/requests/{id}, or send a new request with a new key. |
Limits of the guarantee
Section titled “Limits of the guarantee”- Zero retention covers what Belarel stores and how calls are routed. What a provider does is governed by its own zero-retention commitments, which the routing requirement relies on.
- Content you send elsewhere — your own logs, your observability tools, another API — is outside its scope.
- The mode is set when the key is created and cannot be changed later.

