Baseten Provider
Baseten is an inference platform for serving frontier, enterprise-grade opensource AI models via their API.
Setup
The Baseten provider is available via the @ai-sdk/baseten module. You can install it with
pnpm add @ai-sdk/baseten
Provider Instance
You can import the default provider instance baseten from @ai-sdk/baseten:
import { baseten } from '@ai-sdk/baseten';If you need a customized setup, you can import createBaseten from @ai-sdk/baseten
and create a provider instance with your settings:
import { createBaseten } from '@ai-sdk/baseten';
const baseten = createBaseten({ apiKey: process.env.BASETEN_API_KEY ?? '',});You can use the following optional settings to customize the Baseten provider instance:
-
baseURL string
Use a different URL prefix for API calls, e.g. to use proxy servers. The default prefix is
https://inference.baseten.co/v1. -
apiKey string
API key that is being sent using the
Authorizationheader. It defaults to theBASETEN_API_KEYenvironment variable. It is recommended you set the environment variable usingexportso you do not need to include the field everytime. You can grab your Baseten API Key here -
modelURL string
Custom model URL for specific models (chat or embeddings). If not provided, the default Model APIs will be used.
-
headers Record<string,string>
Custom headers to include in the requests.
-
performanceClient PerformanceClient constructor
Opt in to Baseten's native performance client for embeddings, for client-side batching and request hedging. Pass the
PerformanceClientconstructor from@basetenlabs/performance-client, which you install yourself. When omitted, embeddings use plain HTTP. See Native performance client. -
fetch (input: RequestInfo, init?: RequestInit) => Promise<Response>
Custom fetch implementation.
Model APIs
You can select Baseten models using a provider instance.
The first argument is the model id, e.g. 'moonshotai/Kimi-K2-Instruct-0905': The complete supported models under Model APIs can be found here.
const model = baseten('moonshotai/Kimi-K2-Instruct-0905');Example
You can use Baseten language models to generate text with the generateText function:
import { baseten } from '@ai-sdk/baseten';import { generateText } from 'ai';
const { text } = await generateText({ model: baseten('moonshotai/Kimi-K2-Instruct-0905'), prompt: 'What is the meaning of life? Answer in one sentence.',});Baseten language models can also be used in the streamText function
(see AI SDK Core).
Dedicated Models
Baseten supports dedicated model URLs for both chat and embedding models. You have to specify a modelURL when creating the provider:
OpenAI-Compatible Endpoints (/sync/v1)
For models deployed with Baseten's OpenAI-compatible endpoints:
import { createBaseten } from '@ai-sdk/baseten';
const baseten = createBaseten({ modelURL: 'https://model-{MODEL_ID}.api.baseten.co/sync/v1',});// No modelId is needed because we specified modelURLconst model = baseten();const { text } = await generateText({ model: model, prompt: 'Say hello from a Baseten chat model!',});/predict Endpoints
/predict endpoints are currently NOT supported for chat models. You must use /sync/v1 endpoints for chat functionality.
Embedding Models
You can create models that call the Baseten embeddings API using the .textEmbeddingModel() factory method. Baseten Embeddings Inference deployments are OpenAI-compatible, so embeddings use plain HTTP by default with no extra dependencies.
Important: Embedding models require a dedicated deployment with a custom
modelURL. Unlike chat models, embeddings cannot use Baseten's default Model
APIs and must specify a dedicated model endpoint.
import { createBaseten } from '@ai-sdk/baseten';import { embed, embedMany } from 'ai';
const baseten = createBaseten({ modelURL: 'https://model-{MODEL_ID}.api.baseten.co/sync',});
const embeddingModel = baseten.textEmbeddingModel();
// Single embeddingconst { embedding } = await embed({ model: embeddingModel, value: 'sunny day at the beach',});
// Batch embeddingsconst { embeddings } = await embedMany({ model: embeddingModel, values: [ 'sunny day at the beach', 'rainy afternoon in the city', 'snowy mountain peak', ],});Each request sends at most 128 values. embedMany splits larger inputs into
chunks of that size and runs them in parallel, so you can pass as many values as
you like.
Endpoint Support for Embeddings
Supported:
/syncendpoints (/v1/embeddingsis appended for you)/sync/v1endpoints
Not Supported:
/predictendpoints
Native performance client (optional)
Baseten also publishes @basetenlabs/performance-client, a native client that
adds client-side batching and request hedging on top of the server-side dynamic
batching your deployment already does. It is not installed by default: it is
a native addon, so it cannot load in edge runtimes and bundlers cannot resolve
its platform binaries.
To use it, install it yourself and pass the constructor:
npm i @basetenlabs/performance-clientimport { createBaseten } from '@ai-sdk/baseten';import { PerformanceClient } from '@basetenlabs/performance-client';
const baseten = createBaseten({ modelURL: 'https://model-{MODEL_ID}.api.baseten.co/environments/production/sync', performanceClient: PerformanceClient,});When you opt in, the client handles batching itself, so values are sent in a single call rather than being split at 128.
Error Handling
The Baseten provider includes built-in error handling for common API errors:
import { baseten } from '@ai-sdk/baseten';import { generateText } from 'ai';
try { const { text } = await generateText({ model: baseten('moonshotai/Kimi-K2-Instruct-0905'), prompt: 'Hello, world!', });} catch (error) { console.error('Baseten API error:', error.message);}Common Error Scenarios
// Embeddings require a modelURLtry { baseten.textEmbeddingModel();} catch (error) { // Error: "No model URL provided for embeddings. Please set modelURL option for embeddings."}
// /predict endpoints are not supported for chat modelstry { const baseten = createBaseten({ modelURL: 'https://model-{MODEL_ID}.api.baseten.co/environments/production/predict', }); baseten(); // This will throw an error} catch (error) { // Error: "Not supported. You must use a /sync/v1 endpoint for chat models."}
// /sync/v1 endpoints are now supported for embeddingsconst baseten = createBaseten({ modelURL: 'https://model-{MODEL_ID}.api.baseten.co/environments/production/sync/v1',});const embeddingModel = baseten.textEmbeddingModel(); // This works fine!
// /predict endpoints are not supported for embeddingstry { const baseten = createBaseten({ modelURL: 'https://model-{MODEL_ID}.api.baseten.co/environments/production/predict', }); baseten.textEmbeddingModel(); // This will throw an error} catch (error) { // Error: "Not supported. You must use a /sync or /sync/v1 endpoint for embeddings."}
// Image models are not supportedtry { baseten.imageModel('test-model');} catch (error) { // Error: NoSuchModelError for imageModel}For more information about Baseten models and deployment options, see the Baseten documentation.