Vercel AI SDK
The Vercel AI SDK talks to Belarel through its OpenAI provider, @ai-sdk/openai. You create a provider instance with Belarel’s base URL and your key, then pick the Chat Completions model factory. Everything else — generateText, streamText, tools, structured output — is standard AI SDK code.
Install
Section titled “Install”npm install ai @ai-sdk/openaiCreate the provider
Section titled “Create the provider”import { createOpenAI } from '@ai-sdk/openai';
export const belarel = createOpenAI({ baseURL: 'https://api.belarel.com/v1', apiKey: process.env.BELAREL_API_KEY, // Optional: name your app in usage reports (recorded as the tag `app`). headers: { 'X-Title': 'Acme Assistant' },});Use .chat(), not the default model factory
Section titled “Use .chat(), not the default model factory”Recent versions of @ai-sdk/openai route belarel('<model-id>') to the OpenAI Responses API. Belarel serves Responses too, but statelessly: it stores nothing and refuses item_reference inputs. When the AI SDK replays a multi-turn conversation built from earlier results, it can send earlier assistant turns as item_reference, which Belarel answers with 400 unsupported_parameter.
Use the Chat Completions factory instead — it always sends the full conversation:
const model = belarel.chat('<model-id>');Model ids come from GET /v1/models for your key. See Models.
generateText
Section titled “generateText”import { generateText } from 'ai';import { belarel } from './belarel';
const result = await generateText({ model: belarel.chat('<model-id>'), prompt: 'Write a one-line product tagline for a bike shop.', maxOutputTokens: 100,});
console.log(result.text);console.log(result.usage);
// The raw JSON answer carries Belarel's block: request id, model served, cost.const body = result.response.body as { belarel?: { request_id: string; cost: unknown } };console.log(body.belarel?.request_id, body.belarel?.cost);streamText
Section titled “streamText”import { streamText } from 'ai';import { belarel } from './belarel';
const result = streamText({ model: belarel.chat('<model-id>'), prompt: 'Explain rate limiting in three sentences.',});
for await (const delta of result.textStream) { process.stdout.write(delta);}console.log('\n', await result.usage);Belarel closes every stream with a final chunk that carries usage and the belarel block. The AI SDK reads the usage for you. To read the cost as well, set includeRawChunks: true and look for the raw chunk that holds belarel:
const result = streamText({ model: belarel.chat('<model-id>'), prompt: 'Explain rate limiting in three sentences.', includeRawChunks: true,});
for await (const part of result.fullStream) { if (part.type === 'text-delta') process.stdout.write(part.text); if (part.type === 'raw') { const chunk = part.rawValue as { belarel?: { cost: unknown } }; if (chunk.belarel) console.log('\ncost:', chunk.belarel.cost); }}Tools and structured output
Section titled “Tools and structured output”Tools you define with tool() and structured output with Output or generateObject map onto Chat Completions tools and response_format, which Belarel reads on models that support them. Check belarel.capabilities.tools and belarel.capabilities.structured_output in /v1/models before you rely on either:
- tools sent to a model without tool support are refused with
400 model_lacks_tools; - structured output on a model without support is refused with
400 model_lacks_structured_output.
Most providerOptions.openai settings have no Belarel equivalent and are refused with 400 unsupported_parameter, whose param names the field as sent. For example: logprobs, logitBias, prediction, promptCacheKey, safetyIdentifier and store: true. textVerbosity depends on the route: .chat() sends it as a top-level verbosity, which is refused; .responses() sends it nested as text.verbosity, which Belarel does not check, so it is ignored without an error. user, parallelToolCalls, serviceTier and metadata are accepted. topP, stopSequences, seed, presencePenalty and frequencyPenalty are forwarded to the model; when its provider drops one, the raw answer’s belarel.ignored_parameters names it.
Long calls
Section titled “Long calls”A call has a little under 300 seconds, and a non-streamed answer sends nothing until it is complete. For reasoning models or long outputs, prefer streamText over generateText: tokens arrive as they are generated, and Belarel keeps the connection alive with a comment every 15 seconds. A generateText call that runs out of time fails with 504 request_timeout. See Streaming.
Embeddings
Section titled “Embeddings”The provider’s embedding model requests float vectors (Belarel also serves base64, the default of the OpenAI SDKs). Use the one embedding model Belarel offers:
import { embedMany } from 'ai';import { belarel } from './belarel';
const { embeddings } = await embedMany({ model: belarel.embedding('openai/text-embedding-3-small'), values: ['sunny day at the beach', 'rainy afternoon in the city'],});Your key needs the embeddings scope. On older AI SDK versions, the factory is named textEmbeddingModel.
Checklist
Section titled “Checklist”- Create the provider with
baseURL: 'https://api.belarel.com/v1'andBELAREL_API_KEY. - Use
belarel.chat(id)— orbelarel.responses(id)withstore: false. - Take model ids from
GET /v1/models. - Develop against a
bel_test_…key: answers come from a deterministic sandbox and are never billed.

