Skip to content

Vercel AI SDK

The Vercel AI SDK talks to Belarel through its OpenAI provider, @ai-sdk/openai. You create a provider instance with Belarel’s base URL and your key, then pick the Chat Completions model factory. Everything else — generateText, streamText, tools, structured output — is standard AI SDK code.

Terminal window
npm install ai @ai-sdk/openai
belarel.ts
import { createOpenAI } from '@ai-sdk/openai';
export const belarel = createOpenAI({
baseURL: 'https://api.belarel.com/v1',
apiKey: process.env.BELAREL_API_KEY,
// Optional: name your app in usage reports (recorded as the tag `app`).
headers: { 'X-Title': 'Acme Assistant' },
});

Use .chat(), not the default model factory

Section titled “Use .chat(), not the default model factory”

Recent versions of @ai-sdk/openai route belarel('<model-id>') to the OpenAI Responses API. Belarel serves Responses too, but statelessly: it stores nothing and refuses item_reference inputs. When the AI SDK replays a multi-turn conversation built from earlier results, it can send earlier assistant turns as item_reference, which Belarel answers with 400 unsupported_parameter.

Use the Chat Completions factory instead — it always sends the full conversation:

const model = belarel.chat('<model-id>');

Model ids come from GET /v1/models for your key. See Models.

import { generateText } from 'ai';
import { belarel } from './belarel';
const result = await generateText({
model: belarel.chat('<model-id>'),
prompt: 'Write a one-line product tagline for a bike shop.',
maxOutputTokens: 100,
});
console.log(result.text);
console.log(result.usage);
// The raw JSON answer carries Belarel's block: request id, model served, cost.
const body = result.response.body as { belarel?: { request_id: string; cost: unknown } };
console.log(body.belarel?.request_id, body.belarel?.cost);
import { streamText } from 'ai';
import { belarel } from './belarel';
const result = streamText({
model: belarel.chat('<model-id>'),
prompt: 'Explain rate limiting in three sentences.',
});
for await (const delta of result.textStream) {
process.stdout.write(delta);
}
console.log('\n', await result.usage);

Belarel closes every stream with a final chunk that carries usage and the belarel block. The AI SDK reads the usage for you. To read the cost as well, set includeRawChunks: true and look for the raw chunk that holds belarel:

const result = streamText({
model: belarel.chat('<model-id>'),
prompt: 'Explain rate limiting in three sentences.',
includeRawChunks: true,
});
for await (const part of result.fullStream) {
if (part.type === 'text-delta') process.stdout.write(part.text);
if (part.type === 'raw') {
const chunk = part.rawValue as { belarel?: { cost: unknown } };
if (chunk.belarel) console.log('\ncost:', chunk.belarel.cost);
}
}

Tools you define with tool() and structured output with Output or generateObject map onto Chat Completions tools and response_format, which Belarel reads on models that support them. Check belarel.capabilities.tools and belarel.capabilities.structured_output in /v1/models before you rely on either:

  • tools sent to a model without tool support are refused with 400 model_lacks_tools;
  • structured output on a model without support is refused with 400 model_lacks_structured_output.

Most providerOptions.openai settings have no Belarel equivalent and are refused with 400 unsupported_parameter, whose param names the field as sent. For example: logprobs, logitBias, prediction, promptCacheKey, safetyIdentifier and store: true. textVerbosity depends on the route: .chat() sends it as a top-level verbosity, which is refused; .responses() sends it nested as text.verbosity, which Belarel does not check, so it is ignored without an error. user, parallelToolCalls, serviceTier and metadata are accepted. topP, stopSequences, seed, presencePenalty and frequencyPenalty are forwarded to the model; when its provider drops one, the raw answer’s belarel.ignored_parameters names it.

A call has a little under 300 seconds, and a non-streamed answer sends nothing until it is complete. For reasoning models or long outputs, prefer streamText over generateText: tokens arrive as they are generated, and Belarel keeps the connection alive with a comment every 15 seconds. A generateText call that runs out of time fails with 504 request_timeout. See Streaming.

The provider’s embedding model requests float vectors (Belarel also serves base64, the default of the OpenAI SDKs). Use the one embedding model Belarel offers:

import { embedMany } from 'ai';
import { belarel } from './belarel';
const { embeddings } = await embedMany({
model: belarel.embedding('openai/text-embedding-3-small'),
values: ['sunny day at the beach', 'rainy afternoon in the city'],
});

Your key needs the embeddings scope. On older AI SDK versions, the factory is named textEmbeddingModel.

  1. Create the provider with baseURL: 'https://api.belarel.com/v1' and BELAREL_API_KEY.
  2. Use belarel.chat(id) — or belarel.responses(id) with store: false.
  3. Take model ids from GET /v1/models.
  4. Develop against a bel_test_… key: answers come from a deterministic sandbox and are never billed.