Skip to content

Migrate from OpenAI

Belarel speaks the OpenAI wire format, so the official OpenAI SDKs work once you change two settings: the base URL and the API key. This page covers that change, then the short list of places where Belarel behaves differently and what to do about each.

Create a key in your organization’s dashboard (bel_test_… for the sandbox, bel_live_… for production), store it in BELAREL_API_KEY, and point your client at https://api.belarel.com/v1.

Terminal window
curl https://api.belarel.com/v1/chat/completions \
-H "Authorization: Bearer $BELAREL_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "<model-id>", "messages": [{"role": "user", "content": "Hello"}], "max_tokens": 64}'
  • POST /v1/chat/completions — messages, images on vision-capable models, streaming, client-side tools (tools → tool_calls), structured outputs (response_format with json_schema or json_object), and reasoning_effort.
  • POST /v1/responses — string or item input, instructions, function tools, text.format, reasoning.effort, and streaming events, as long as you send the whole conversation each time (see below).
  • POST /v1/embeddings — with one model (see below). The SDKs’ default base64 encoding works as is.
  • The error shape { "error": { "message", "type", "code", "param" } }, so your SDK raises the same exception classes.

Model ids are the ones GET /v1/models returns for your key — the models your organization enabled, priced at your organization’s rates. A model your key cannot use answers 404 model_not_found; one your organization’s policy blocks answers 403 model_not_allowed. Read the list instead of hard-coding OpenAI’s names. See Models.

Nothing is stored on Belarel’s side. store: true, previous_response_id, and item_reference input items are refused with 400 unsupported_parameter; send the full conversation on every call, as you would with Chat Completions. input_file is not supported yet, tool_choice accepts only auto and none, and include, truncation: "auto" and background: true are refused.

Belarel serves one embedding model, openai/text-embedding-3-small, at 1536 dimensions. Any other model is a 400, never a silent substitution. encoding_format can be float or base64, so the SDKs work without a setting:

client.embeddings.create(model="openai/text-embedding-3-small", input=["hello"])

Unsupported parameters are refused, not ignored

Section titled “Unsupported parameters are refused, not ignored”

Like OpenAI, Belarel refuses a top-level field it does not support with 400 unsupported_parameter, and param names the field. This applies to Chat Completions and Responses; on those two endpoints, no top-level field is dropped without telling you.

  • Forwarded to the model: temperature, top_p and tool_choice (auto or none) on both endpoints; stop (up to 4), seed, presence_penalty and frequency_penalty on Chat Completions only — Responses refuses them. Plus tools, response_format (text.format on Responses) and reasoning_effort (reasoning.effort) on models that support them. Each model’s supported_parameters in /v1/models lists what is forwarded to it, in Chat Completions terms: it includes stop, seed and the penalties even though Responses refuses them.
  • Accepted with no effect: user, parallel_tool_calls and service_tier on both endpoints; stream_options on Chat Completions only — Responses refuses it.
  • Refused by name: n above 1, logprobs, top_logprobs, logit_bias, store: true, prediction, audio and non-text modalities, a tool_choice that forces a tool, the legacy functions, and any top-level field Belarel does not know, such as web_search_options.

Two limits to know. Fields nested inside an accepted object — such as text.verbosity or reasoning.summary on Responses, or json_schema.strict — are not checked: one Belarel does not use is ignored without an error. And /v1/embeddings does not refuse unknown fields: it ignores them.

A provider can still drop a forwarded parameter, for example seed on some models. The call succeeds, and belarel.ignored_parameters lists what was dropped. See Chat Completions.

Three related behaviors:

  • tools sent to a model whose belarel.capabilities.tools is false are refused with 400 model_lacks_tools.
  • finish_reason is the model’s, as with OpenAI: stop, length, content_filter or tool_calls.
  • Chat Completions does not accept the developer role (400 invalid_messages). Use system.

A call has a little under 300 seconds. A non-streamed call that runs longer gets 504 request_timeout, with x-should-retry: false so your SDK does not replay it. Use stream: true for reasoning models and long outputs. See Streaming.

Every key has a max_output_tokens ceiling. A max_tokens above it is refused with 400 max_tokens_exceeds_key_limit; when you omit max_tokens, the key’s ceiling applies.

JSON answers include a belarel object next to usage: the request id, the model that served the call, the cost in USD and credits, the data mode, and the policy decisions. When you stream, the last chunk before [DONE] has an empty choices array and carries usage and belarel — you don’t need stream_options. If the cost isn’t known in time, cost.source is pending; read it later with GET /v1/requests/{id}. See Cost and usage.

The metadata object is read as tags that label the call in your usage reports. Values must be strings, at most 16 tags per call; a malformed tag is a 400 invalid_tags. See Tags and attribution.

The shape is OpenAI’s, and so is 401 invalid_api_key for a missing, unknown or malformed key. Belarel adds codes: 402 for money (key_cap_exceeded, insufficient_credits, budget_exceeded, max_cost_per_request, spend_refused), 429 with Retry-After for rate_limited and too_many_in_flight, and 503 when Belarel could not decide — retry those, except two that a retry will not fix: 503 pricing_unavailable, and a 503 admission_unavailable whose message says BYOK is not priced for your organization. That BYOK answer still carries Retry-After: 60, so the official SDKs retry it twice on their own before giving up; lower maxRetries (max_retries in Python), or stop on that message in your own retry logic. Policy refusals list the rule that refused in belarel.policy.decisions. See Errors.

  1. Set the base URL to https://api.belarel.com/v1 and the key to BELAREL_API_KEY.
  2. Replace hard-coded model names with ids from GET /v1/models.
  3. Remove store, previous_response_id, and item_reference from Responses calls; send full history.
  4. Remove the parameters Belarel refuses (n > 1, logprobs, logit_bias, a forced tool_choice, functions…). A 400 unsupported_parameter names each one.
  5. Check that your max_tokens values fit under your key’s max_output_tokens.
  6. Turn on stream: true for calls that can run for minutes.
  7. Handle 402 and 403 as final, 429 and 503 as retryable — except 503 pricing_unavailable and a BYOK 503 admission_unavailable — and 504 by streaming.
  8. Run your test suite with a bel_test_… key first — it is never billed.