Migrate from OpenAI
Belarel speaks the OpenAI wire format, so the official OpenAI SDKs work once you change two settings: the base URL and the API key. This page covers that change, then the short list of places where Belarel behaves differently and what to do about each.
Change the base URL and the key
Section titled “Change the base URL and the key”Create a key in your organization’s dashboard (bel_test_… for the sandbox, bel_live_… for production), store it in BELAREL_API_KEY, and point your client at https://api.belarel.com/v1.
curl https://api.belarel.com/v1/chat/completions \ -H "Authorization: Bearer $BELAREL_API_KEY" \ -H "Content-Type: application/json" \ -d '{"model": "<model-id>", "messages": [{"role": "user", "content": "Hello"}], "max_tokens": 64}'import osfrom openai import OpenAI
client = OpenAI(base_url="https://api.belarel.com/v1", api_key=os.environ["BELAREL_API_KEY"])
r = client.chat.completions.create( model="<model-id>", messages=[{"role": "user", "content": "Hello"}], max_tokens=64,)print(r.choices[0].message.content)print(r.model_dump()["belarel"]["cost"])import OpenAI from 'openai';
const client = new OpenAI({ baseURL: 'https://api.belarel.com/v1', apiKey: process.env.BELAREL_API_KEY });
const r = await client.chat.completions.create({ model: '<model-id>', messages: [{ role: 'user', content: 'Hello' }], max_tokens: 64,});console.log(r.choices[0].message.content);console.log((r as unknown as { belarel: { cost: unknown } }).belarel.cost);What works unchanged
Section titled “What works unchanged”- POST
/v1/chat/completions— messages, images on vision-capable models, streaming, client-side tools (tools→tool_calls), structured outputs (response_formatwithjson_schemaorjson_object), andreasoning_effort. - POST
/v1/responses— string or iteminput,instructions, function tools,text.format,reasoning.effort, and streaming events, as long as you send the whole conversation each time (see below). - POST
/v1/embeddings— with one model (see below). The SDKs’ defaultbase64encoding works as is. - The error shape
{ "error": { "message", "type", "code", "param" } }, so your SDK raises the same exception classes.
What differs
Section titled “What differs”Model ids come from your catalog
Section titled “Model ids come from your catalog”Model ids are the ones GET /v1/models returns for your key — the models your organization enabled, priced at your organization’s rates. A model your key cannot use answers 404 model_not_found; one your organization’s policy blocks answers 403 model_not_allowed. Read the list instead of hard-coding OpenAI’s names. See Models.
Responses is stateless
Section titled “Responses is stateless”Nothing is stored on Belarel’s side. store: true, previous_response_id, and item_reference input items are refused with 400 unsupported_parameter; send the full conversation on every call, as you would with Chat Completions. input_file is not supported yet, tool_choice accepts only auto and none, and include, truncation: "auto" and background: true are refused.
Embeddings: one model
Section titled “Embeddings: one model”Belarel serves one embedding model, openai/text-embedding-3-small, at 1536 dimensions. Any other model is a 400, never a silent substitution. encoding_format can be float or base64, so the SDKs work without a setting:
client.embeddings.create(model="openai/text-embedding-3-small", input=["hello"])Unsupported parameters are refused, not ignored
Section titled “Unsupported parameters are refused, not ignored”Like OpenAI, Belarel refuses a top-level field it does not support with 400 unsupported_parameter, and param names the field. This applies to Chat Completions and Responses; on those two endpoints, no top-level field is dropped without telling you.
- Forwarded to the model:
temperature,top_pandtool_choice(autoornone) on both endpoints;stop(up to 4),seed,presence_penaltyandfrequency_penaltyon Chat Completions only — Responses refuses them. Plustools,response_format(text.formaton Responses) andreasoning_effort(reasoning.effort) on models that support them. Each model’ssupported_parametersin/v1/modelslists what is forwarded to it, in Chat Completions terms: it includesstop,seedand the penalties even though Responses refuses them. - Accepted with no effect:
user,parallel_tool_callsandservice_tieron both endpoints;stream_optionson Chat Completions only — Responses refuses it. - Refused by name:
nabove 1,logprobs,top_logprobs,logit_bias,store: true,prediction,audioand non-textmodalities, atool_choicethat forces a tool, the legacyfunctions, and any top-level field Belarel does not know, such asweb_search_options.
Two limits to know. Fields nested inside an accepted object — such as text.verbosity or reasoning.summary on Responses, or json_schema.strict — are not checked: one Belarel does not use is ignored without an error. And /v1/embeddings does not refuse unknown fields: it ignores them.
A provider can still drop a forwarded parameter, for example seed on some models. The call succeeds, and belarel.ignored_parameters lists what was dropped. See Chat Completions.
Three related behaviors:
toolssent to a model whosebelarel.capabilities.toolsisfalseare refused with400 model_lacks_tools.finish_reasonis the model’s, as with OpenAI:stop,length,content_filterortool_calls.- Chat Completions does not accept the
developerrole (400 invalid_messages). Usesystem.
Long calls should stream
Section titled “Long calls should stream”A call has a little under 300 seconds. A non-streamed call that runs longer gets 504 request_timeout, with x-should-retry: false so your SDK does not replay it. Use stream: true for reasoning models and long outputs. See Streaming.
Output length is bounded by your key
Section titled “Output length is bounded by your key”Every key has a max_output_tokens ceiling. A max_tokens above it is refused with 400 max_tokens_exceeds_key_limit; when you omit max_tokens, the key’s ceiling applies.
Every answer carries a belarel block
Section titled “Every answer carries a belarel block”JSON answers include a belarel object next to usage: the request id, the model that served the call, the cost in USD and credits, the data mode, and the policy decisions. When you stream, the last chunk before [DONE] has an empty choices array and carries usage and belarel — you don’t need stream_options. If the cost isn’t known in time, cost.source is pending; read it later with GET /v1/requests/{id}. See Cost and usage.
metadata becomes attribution tags
Section titled “metadata becomes attribution tags”The metadata object is read as tags that label the call in your usage reports. Values must be strings, at most 16 tags per call; a malformed tag is a 400 invalid_tags. See Tags and attribution.
More error codes
Section titled “More error codes”The shape is OpenAI’s, and so is 401 invalid_api_key for a missing, unknown or malformed key. Belarel adds codes: 402 for money (key_cap_exceeded, insufficient_credits, budget_exceeded, max_cost_per_request, spend_refused), 429 with Retry-After for rate_limited and too_many_in_flight, and 503 when Belarel could not decide — retry those, except two that a retry will not fix: 503 pricing_unavailable, and a 503 admission_unavailable whose message says BYOK is not priced for your organization. That BYOK answer still carries Retry-After: 60, so the official SDKs retry it twice on their own before giving up; lower maxRetries (max_retries in Python), or stop on that message in your own retry logic. Policy refusals list the rule that refused in belarel.policy.decisions. See Errors.
Migration checklist
Section titled “Migration checklist”- Set the base URL to
https://api.belarel.com/v1and the key toBELAREL_API_KEY. - Replace hard-coded model names with ids from
GET /v1/models. - Remove
store,previous_response_id, anditem_referencefrom Responses calls; send full history. - Remove the parameters Belarel refuses (
n> 1,logprobs,logit_bias, a forcedtool_choice,functions…). A400 unsupported_parameternames each one. - Check that your
max_tokensvalues fit under your key’smax_output_tokens. - Turn on
stream: truefor calls that can run for minutes. - Handle
402and403as final,429and503as retryable — except503 pricing_unavailableand a BYOK503 admission_unavailable— and504by streaming. - Run your test suite with a
bel_test_…key first — it is never billed.

