Skip to content

Sandbox & testing

A sandbox key (bel_test_…) runs your integration end to end without reaching a model provider and without spending anything. Use it in development and in CI; switch to a live key (bel_live_…) for real traffic.

An owner or administrator creates it in the organization dashboard, under Governance → API keys, with Mode set to Sandbox. Sub-keys created by a sandbox key are sandbox keys too.

Sandbox calls are served by a mock provider that never touches the network. It is deterministic: the same request always produces the same output.

Chat Completions and Responses

  • The answer is your last user message, prefixed with [sandbox] (the echo is cut at 200 characters). For example, "Count to three." answers "[sandbox] Count to three.".
  • If you declare function tools, the first call returns a call to the first tool you declared, with arguments {}, and finish_reason: "tool_calls". Once you send the tool result back, the next answer is text and the turn ends — the sandbox never loops.
  • Token usage is estimated from characters, about four characters per token.
  • Streaming works the same way as live, including the closing chunk with usage and belarel.

Embeddings

  • Each input gets a 1536-dimension, unit-length vector computed from its text: the same text always gets the same vector, so your indexing and search code runs for real.

Every answer reports the sandbox in the cost of its belarel block:

"cost": { "usd": 0, "credits": 0, "source": "estimated", "billing": "sandbox" }

In GET /v1/models, the list returned to a sandbox key carries belarel.sandbox: true.

Faithful to live Different from live
Authentication, scopes and the uniform 401 The content of answers (always the echo)
Request validation and error shapes, including 400 unsupported_parameter and 400 model_lacks_tools Latency, and provider failures
Your key’s model catalog and policy refusals (404 model_not_found, 403 model_not_allowed) Token counts (estimated from characters)
Required tags, default tags, X-Belarel-Tags Cost: always zero, never billed
Requests per minute, tokens per minute, calls in flight Spend caps: a sandbox call reserves nothing, so 402 key_cap_exceeded never happens
The key’s maximum output tokens Your organization’s credits, budgets and contract ceiling are not checked
X-Request-Id, the belarel block, idempotency model_not_zdr and model_not_priced are not raised
Streaming format and tool-call relay Reasoning content and web search results
finish_reason is always stop or tool_calls: the echo is never cut by max_tokens, and belarel.ignored_parameters never appears

Most refusals can be produced on purpose with a sandbox key or one of its sub-keys:

To see Do this
401 invalid_api_key Send a token that does not exist, or a revoked key.
403 scope_denied Create a sub-key with "scopes": ["usage:read"] and call POST /v1/chat/completions.
404 model_not_found Call a model id that is not in your catalog.
403 model_not_allowed Create a sub-key with allowed_model_keys and call another model of your catalog.
400 missing_tags Create a sub-key with "required_tags": ["customer"] and call it without that tag.
400 max_tokens_exceeds_key_limit Send a max_tokens above the key’s maximum output tokens.
429 rate_limited Create a sub-key with "caps": { "rpm": 1 } and send two calls in the same minute. Check Retry-After.
422 idempotency_key_reused Send two different bodies with the same Idempotency-Key.
Tool calling Declare a function tool and complete the tool round trip.

Sub-key creation is described in Authentication & keys.

For spend refusals (402 key_cap_exceeded and the rest), use a live key with a very small daily cap, for example a live sub-key with "caps": { "usd_day": 0.01 }. Those calls are billed.

Create a live key and swap BELAREL_API_KEY. Your code does not change. Before you do, check that your integration:

  • Handles 402 refusals (key caps, credits, budgets, and the cost ceiling per request).
  • Retries 429 and 503 after Retry-After, and does not retry other 4xx.
  • Accepts belarel.cost.source: "pending" and reads the settled cost later from GET /v1/requests/{id} if it needs it.
  • Sets max_tokens deliberately: it decides how much each call reserves against your caps.
  • Uses one key per service and environment, with caps set.