Skip to content

Embeddings

POST /v1/embeddings turns text into vectors for search, clustering and retrieval. It speaks the OpenAI Embeddings dialect and serves one model: openai/text-embedding-3-small, which returns vectors of 1536 dimensions.

The key needs the embeddings scope. A key with only inference gets 403 scope_denied.

Terminal window
curl https://api.belarel.com/v1/embeddings \
-H "Authorization: Bearer $BELAREL_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "openai/text-embedding-3-small",
"input": ["How do I reset my password?", "Where is my invoice?"]
}'

The official OpenAI SDKs need no extra setting: they ask for base64 vectors by default and decode them into lists of floats for you, as they do with OpenAI.

Field Rule
model Required. Must be openai/text-embedding-3-small.
input Required. A non-empty string, or an array of 1 to 256 non-empty strings. Each string is at most 8,000 characters. Arrays of token ids are not accepted.
encoding_format Optional. "float" (the default) returns arrays of numbers. "base64" returns each vector as a base64 string of little-endian 32-bit floats, as OpenAI does. Any other value returns 400 invalid_request.
dimensions Optional. Only 1536. Shortened vectors are not available.
metadata, belarel_tags Optional attribution tags. See Tags and attribution.

Unlike Chat Completions and Responses, this endpoint does not refuse fields it does not know: any other field, such as user, is ignored.

{
"object": "list",
"data": [
{ "object": "embedding", "index": 0, "embedding": [0.0123, -0.0456, "…"] },
{ "object": "embedding", "index": 1, "embedding": [0.0078, 0.0311, "…"] }
],
"model": "openai/text-embedding-3-small",
"usage": { "prompt_tokens": 14, "total_tokens": 14 },
"belarel": {
"request_id": "req_0f1e…",
"model_requested": "openai/text-embedding-3-small",
"model_served": "openai/text-embedding-3-small",
"cost": { "usd": 0.0000003, "credits": 0.0003, "source": "estimated", "billing": "platform" },
"policy": { "decisions": [{ "rule": "embedding_model", "verdict": "allow", "detail": "residency any" }] },
"data_mode": "standard"
}
}

data keeps the order of input: index 0 is your first string. The belarel block has the same fields as in Chat Completions. If the price of the call is not known when the answer is sent, cost.source is pending and usd is null; read the final figure later with request inspection.

To decode a base64 vector yourself, read it as little-endian float32. In Python: numpy.frombuffer(base64.b64decode(s), dtype="<f4").

  • Sandbox keys (bel_test_…) return deterministic vectors: the same text always gives the same unit-length vector of 1536 dimensions. No provider is called and the cost is 0. Use them to test your indexing code, not to judge search quality. See Sandbox.
  • Zero-retention keys can call embeddings. See Zero data retention.
  • BYOK keys cannot call embeddings: embeddings always run on the platform’s provider account. A live BYOK key gets 400 byok_unavailable. Use a key on platform billing for embeddings. See BYOK.
  • Confidential keys get 501 not_implemented. Confidential inference is coming soon.

The model must also be open to your organization’s API and allowed for the key. It appears in GET /v1/models with belarel.category: "embedding" when the key may call it.

Code HTTP Meaning What to do
model_not_allowed 400 You named a model other than openai/text-embedding-3-small. Use that model.
invalid_request 400 Bad input, encoding_format or dimensions. param names the field. Fix the field.
byok_unavailable 400 The key is a BYOK key. Use a key on platform billing.
scope_denied 403 The key lacks the embeddings scope. Use a key with that scope.
model_not_allowed 403 The model is not in this key’s list of allowed models, or your organization’s choices, your contract or your residency zone exclude it. belarel.policy.decisions names the rule. Add it to the key or use another key; for an organization, contract or residency rule, ask your organization admin.
model_not_found 404 The model is not open to your organization’s API. Ask your organization admin.
max_cost_per_request 402 The call would exceed your contract’s cost ceiling per request. Send fewer inputs per call.

Every other code (caps, credits, rate limits, 503) works as on the other endpoints. See Errors. One difference: a failed embeddings call counts nothing against your key’s caps — its reservation is always released.

Embeddings accept an Idempotency-Key, with the same rules as chat. See Idempotency.