Idempotency
A network error leaves you unsure whether your request ran. Retrying blindly can run it, and bill it, twice. Send an Idempotency-Key header and Belarel runs the request at most once per key: a retry of a finished call returns the stored answer instead of calling the model again.
Idempotency works on POST /v1/chat/completions, POST /v1/responses and POST /v1/embeddings.
How it works
Section titled “How it works”- The value. Any string of 1 to 255 characters, after trimming spaces. A UUID per logical operation is a good choice. An empty value, or one longer than 255 characters, is ignored, and the call is then not idempotent.
- The scope. Keys are scoped to the API key that sends them. The same value sent with two different API keys names two different requests.
- The lifetime. 24 hours from the first use. After that, the same value starts a new request.
- The body. The key is bound to the request body. A retry must send the same JSON body: same fields, same values, in the same order. Belarel does not keep your body to compare it: it keeps a keyed fingerprint (HMAC-SHA256 with a server-side secret), which cannot be reversed or confirmed by guessing without that secret. This holds for zero-retention keys too.
- When it is checked. After your key is admitted and the request is validated, and before any money is reserved. A request refused for a validation error (a
400) does not use up the key.
What a retry gets
Section titled “What a retry gets”| Situation | Answer |
|---|---|
| First use | The request runs normally. |
| The first call finished with a JSON answer | 200 with the stored answer, the header Idempotent-Replayed: true, and the original X-Request-Id. The model is not called and nothing is billed again. |
| The first call is still running | 409 idempotency_conflict. The message names the original request (… in progress (request req_…)). |
| The first call was streamed | 409 idempotency_conflict (… already served (request req_…)). Streamed answers are never stored, so they cannot be replayed. |
| The key is a zero-retention key | 409 idempotency_conflict. The answer is never stored, so it cannot be replayed. See Zero data retention. |
| Same key, different body | 422 idempotency_key_reused, param: "Idempotency-Key". |
| Same key, first used before an update to how Belarel fingerprints requests | 422 idempotency_key_reused, even with the same body, until that key’s 24 hours are over. Use a new key. |
| The first call was refused, or failed, before a JSON answer (caps, credits, rate limit, an unreachable model, a timeout) | The key is released. A retry runs as a new request, and is counted as one. See What a failed call costs. |
| Belarel cannot check the key right now | 503 admission_unavailable with Retry-After. The call is refused rather than run without protection. |
On POST /v1/embeddings, the 409 message always says in progress, even when the first call has finished and its answer was not stored.
If the process serving the first call dies before it finishes, the key frees itself after 5 minutes. A retry after that runs as a new request.
Send the header
Section titled “Send the header”curl https://api.belarel.com/v1/chat/completions \ -H "Authorization: Bearer $BELAREL_API_KEY" \ -H "Content-Type: application/json" \ -H "Idempotency-Key: 6f1c2b9e-1d8a-4f7e-9a51-3c0d2e7b8a44" \ -d '{"model": "openai/gpt-5", "messages": [{"role": "user", "content": "Summarize ticket 1182."}]}'import osimport uuidfrom openai import OpenAI
client = OpenAI(base_url="https://api.belarel.com/v1", api_key=os.environ["BELAREL_API_KEY"])
completion = client.chat.completions.create( model="openai/gpt-5", messages=[{"role": "user", "content": "Summarize ticket 1182."}], extra_headers={"Idempotency-Key": str(uuid.uuid4())},)import OpenAI from 'openai';import { randomUUID } from 'node:crypto';
const client = new OpenAI({ baseURL: 'https://api.belarel.com/v1', apiKey: process.env.BELAREL_API_KEY });
const completion = await client.chat.completions.create( { model: 'openai/gpt-5', messages: [{ role: 'user', content: 'Summarize ticket 1182.' }] }, { headers: { 'Idempotency-Key': randomUUID() } },);The official OpenAI SDKs retry 408, 409, 429 and 5xx answers on their own (twice by default), and they resend your headers. With an Idempotency-Key set, those automatic retries never run a call that already succeeded. A call that failed with a 5xx is not stored, so its retry runs again. 504 request_timeout is the exception: it carries x-should-retry: false, so the SDKs do not replay a call that would only time out again. Retry it yourself with stream: true and a new key, since the body has changed. See Streaming.
What a failed call costs
Section titled “What a failed call costs”A retry runs as a new request, so it helps to know what the failed attempt already counted against your key’s spend caps. On Chat Completions and Responses:
| The attempt | What it counts |
|---|---|
| Refused by Belarel before the model ran (validation, policy, caps, credits, rate limits) | Nothing. |
Refused by the provider before any work (for example a 400 for a bad request, or a 429), with nothing generated |
Nothing. |
Failed on the provider’s side (503 service_unavailable) or a stream that died, with nothing generated and nothing metered |
Only the input part of its estimate, at least one token. |
Stopped by the time limit (504 request_timeout) |
What the provider metered before it was stopped or, when nothing was metered, the full estimate. A long reasoning may have run. |
| Anything else: output already produced, or another provider error | What the provider metered or, when nothing was metered, the full estimate. |
The estimate is the amount Belarel reserves before the call, from the size of your input and your output budget, at your organization’s prices. A smaller max_tokens means a smaller estimate.
What the provider metered is billed like any call. An amount settled without metered usage counts only against the key’s spend caps (and its parent’s) — it is not billed: no credits are debited, and it adds nothing to the cost in /v1/usage or /v1/credits. You see it as cost.settled_usd on GET /v1/requests/{id} and in the key’s spend on GET /v1/key. A failed embeddings call counts nothing: its reservation is always released.
A safe retry recipe
Section titled “A safe retry recipe”-
Create the key once per operation, before the first attempt, and keep it with the operation (for example, next to the ticket you are summarizing). Never create a new key inside the retry loop.
-
Send the same body on every attempt. Build it once and reuse it.
-
Retry with the same key on a network error, on
429and on503. Wait forRetry-Afterwhen it is present. Use exponential backoff otherwise. Two503answers will not change on retry:pricing_unavailable, and anadmission_unavailablewhose message says BYOK is not priced for your organization — see Errors. On504 request_timeout, do not retry the same way: send the call withstream: true, under a new key. -
On
409 idempotency_conflict, the first call is still running, or it finished with an answer that was not stored (streamed, or a zero-retention key). Wait and retry: a running call will soon replay its answer. If the conflict persists, the answer cannot be replayed. Look up the request named in the message, and start over with a new key if you still need an answer. -
Stop on any other
4xx.422 idempotency_key_reusedmeans your code changed the body under the same key, which is a bug to fix.400,402,403and404will not change on retry.
import timeimport uuidimport openai
def summarize(client, ticket_text: str, attempts: int = 5): key = str(uuid.uuid4()) # one key for the whole operation body = dict(model="openai/gpt-5", messages=[{"role": "user", "content": ticket_text}]) for attempt in range(attempts): try: return client.with_options(max_retries=0).chat.completions.create( **body, extra_headers={"Idempotency-Key": key} ) except openai.APIConnectionError: pass # network error: retry with the same key except openai.APIStatusError as err: if err.status_code not in (409, 429, 503): raise # 400, 402, 403, 404, 422, 504: retrying as is will not help delay = err.response.headers.get("retry-after") time.sleep(float(delay) if delay else 2 ** attempt) continue time.sleep(2 ** attempt) raise RuntimeError("gave up after retries")
