Use Grok from TypeScript with a typed, ESM client built on the SpaceXAI REST API. The SDK has no runtime dependencies and includes streaming, structured output, function tools, image input, image and video generation, file uploads, batch processing, text to speech and transcription, multi-turn conversations, and access to usage and HTTP metadata.
Experimental. This SDK is in early development. It currently covers the Responses API, image and video generation, the Files, Batch, and Voice APIs, tokenization, and model and account lookup, and its interfaces may change between releases before 1.0. Pin an exact version and read the changelog when upgrading. Feedback and bug reports are welcome in issues.
- Node.js 22.13 or later
- A SpaceXAI API key
- An ESM project
Install the package with your preferred package manager:
npm install @xai-official/sdk
pnpm add @xai-official/sdk
Set your API key in the environment. The client reads XAI_API_KEY automatically.
export XAI_API_KEY="your-api-key"
js
import { SpaceXAI } from "@xai-official/sdk";
const client = new SpaceXAI();
const response = await client.responses.create({
model: "grok-4.7",
input: "Explain why the sky is blue in one sentence.",
});
console.log(response.toText());
Keep API keys on the server. The SDK blocks browser and Worker use by default because shipping a secret key to client-side code exposes it to users.
Set stream: true to receive output as it's generated. Listen for answer text with on("text"), then await stream.done() for the final response:
const stream = await client.responses.create({
model: "grok-4.7",
input: "Write a short story about a curious robot.",
stream: true,
});
const response = await stream
.on("text", (text) => process.stdout.write(text))
.done();
console.log(`\n${response.usage.total_tokens} tokens`);
done() resolves to the same response object a non-streamed request returns. It rejects if the stream fails or closes before the response completes.
Besides "text", on() has helper events for the rest of a response:
"reasoning"for each chunk of reasoning text or reasoning summary"tool_call"for each tool call once its arguments are complete, whether your code or SpaceXAI runs it. Checkcall.typeto tell them apart"client_tool_call"for each call that your code runs: your function tools and shell commands"server_tool_call"for each call to a tool that SpaceXAI runs, such as web search or code execution"image"for each finished image generation call, with the base64 image inresult"citation"for each URL citation in the answer
Each tool call fires once, when its arguments are complete. For your function tools, that's when to run the function, since nothing has run it yet. To show that a call has started before its arguments arrive, listen for "response.output_item.added".
const response = await stream
.on("reasoning", (text) => process.stderr.write(text))
.on("text", (text) => process.stdout.write(text))
.on("tool_call", (call) => console.error(`\n${call.type} ${call.status}`))
.done();
on() also takes any server-sent event type, such as "response.completed", and passes the listener the typed event. Events the SDK doesn't recognize arrive as "unknown", with the original payload in event.raw.
You can also iterate over the stream to handle events in a loop. After the loop, done() resolves right away with the final response:
for await (const event of stream) {
if (event.type === "response.function_call_arguments.delta") {
process.stdout.write(event.delta);
}
}
const response = await stream.done();
console.log(response.toText());
stream.http has the HTTP status, headers, and request IDs as soon as create() returns.
If you stop consuming a stream early, call await stream.close() to cancel its response body.
Responses are not stored by default. Use toInput() to carry the model output, including encrypted reasoning content, into the next turn. Reuse one prompt_cache_key across the conversation to improve prompt cache routing:
import { randomUUID } from "node:crypto";
import { type InputItem, SpaceXAI } from "@xai-official/sdk";
const client = new SpaceXAI();
const promptCacheKey = `conversation:${randomUUID()}`;
const input: Array<InputItem> = [
{ role: "user", content: "My name is Ada. Remember it." },
];
const first = await client.responses.create({
model: "grok-4.7",
input,
prompt_cache_key: promptCacheKey,
});
const second = await client.responses.create({
model: "grok-4.7",
input: [
...input,
...first.toInput(),
{ role: "user", content: "What is my name?" },
],
prompt_cache_key: promptCacheKey,
});
console.log(second.toText());
Use a different cache key for each unrelated conversation.
To continue a stored response by ID, opt in to storage:
const first = await client.responses.create({
model: "grok-4.7",
input: "My name is Ada. Remember it.",
store: true,
});
const second = await client.responses.create({
model: "grok-4.7",
input: "What is my name?",
previous_response_id: first.id,
store: true,
});
Every turn resends the whole conversation, so input tokens grow as it gets longer. Compact the conversation into a single encrypted item with responses.compact(), then start the next input with the compacted output. Continuing the toInput() example:
const compacted = await client.responses.compact({
model: "grok-4.7",
input: [
...input,
...first.toInput(),
{ role: "user", content: "What is my name?" },
...second.toInput(),
],
});
const third = await client.responses.create({
model: "grok-4.7",
input: [
...compacted.output,
{ role: "user", content: "Spell my name backwards." },
],
prompt_cache_key: promptCacheKey,
});
console.log(third.toText());
Pass compacted.output unchanged and add new turns after it. The conversation must still fit in the model's context window when you compact it. compacted.usage reports the tokens the compaction used and dropped_message_count, the number of messages it replaced.
Pass an image URL alongside text:
const response = await client.responses.create({
model: "grok-4.7",
input: [
{
role: "user",
content: [
{ type: "input_text", text: "Describe this image." },
{
type: "input_image",
image_url: "https://example.com/image.jpg",
detail: "high",
},
],
},
],
});
console.log(response.toText());
For a local image, pass a Blob or File as image. The SDK converts it to a data URL before sending the request, and detects JPEG, PNG, or WebP from the bytes when the Blob has no MIME type.
import { readFile } from "node:fs/promises";
const bytes = new Uint8Array(await readFile("./image.png"));
const image = new Blob([bytes], { type: "image/png" });
const response = await client.responses.create({
model: "grok-4.7",
input: [
{
role: "user",
content: [
{ type: "input_text", text: "What is in this image?" },
{ type: "input_image", image },
],
},
],
});
Provide a JSON Schema through text.format. Call toJson() to parse the completed text output.
const response = await client.responses.create({
model: "grok-4.7",
input: "Give me a city to visit in Japan.",
text: {
format: {
type: "json_schema",
name: "travel_suggestion",
schema: {
type: "object",
properties: {
city: { type: "string" },
reason: { type: "string" },
},
required: ["city", "reason"],
additionalProperties: false,
},
},
},
});
const suggestion = response.toJson();
console.log(suggestion);
toJson() returns unknown. Validate the result before using it at a trust boundary. The non-throwing response.parsed getter returns null when the output is incomplete or is not valid JSON.
Describe functions with JSON Schema, run the requested function in your application, then return its output to the model:
import { type InputItem, type Tool, isFunctionCall, SpaceXAI } from "@xai-official/sdk";
const client = new SpaceXAI();
const prompt = "What is the weather in San Francisco?";
const input: Array<InputItem> = [{ role: "user", content: prompt }];
const getWeatherTool: Tool = {
type: "function",
name: "get_weather",
description: "Get the current weather for a location.",
parameters: {
type: "object",
properties: {
location: { type: "string" },
},
required: ["location"],
additionalProperties: false,
},
};
async function getWeather(args: unknown) {
if (
typeof args !== "object" ||
args === null ||
!("location" in args) ||
typeof args.location !== "string"
) {
throw new Error("get_weather expects a location string");
}
return { location: args.location, temperatureC: 18, conditions: "sunny" };
}
const response = await client.responses.create({
model: "grok-4.7",
input,
parallel_tool_calls: false,
tools: [getWeatherTool],
});
const call = response.output.find(isFunctionCall);
if (!call) throw new Error("The model did not call get_weather");
const result = await getWeather(JSON.parse(call.arguments));
const answer = await client.responses.create({
model: "grok-4.7",
input: [
...input,
...response.toInput(),
{ type: "function_call_output", call_id: call.call_id, output: JSON.stringify(result) },
],
});
console.log(answer.toText());
parallel_tool_calls: false limits the model to one function call per turn, so this example only has to handle one. By default the model can ask for several at once, which the tool call loop handles.
Treat function names and arguments as untrusted input. Only dispatch functions you have explicitly allowed, and validate arguments before executing them, as getWeather() does.
To let the model call tools until it has an answer, run a loop. Stream a turn, run each function call as soon as it finishes streaming, then send the outputs back along with the model's output. Stop when a turn makes no function calls, and cap the number of turns so a model that keeps calling tools can't loop forever. This reuses getWeatherTool and getWeather from the example above:
import { type FunctionToolCall, type InputItem, type Tool, SpaceXAI } from "@xai-official/sdk";
const client = new SpaceXAI();
const tools: Array<Tool> = [getWeatherTool];
const handlers: Record<string, (args: unknown) => Promise<unknown>> = {
get_weather: getWeather,
};
async function runTool(call: FunctionToolCall): Promise<InputItem> {
let output: unknown;
try {
const handler = handlers[call.name];
if (!handler) throw new Error(`Unknown tool: ${call.name}`);
output = await handler(JSON.parse(call.arguments));
} catch (err) {
output = { error: err instanceof Error ? err.message : String(err) };
}
return {
type: "function_call_output",
call_id: call.call_id,
output: JSON.stringify(output),
};
}
const input: Array<InputItem> = [
{ role: "user", content: "Compare the weather in Paris and Tokyo." },
];
for (let turn = 0; turn < 10; turn++) {
const toolRuns: Array<Promise<InputItem>> = [];
const stream = await client.responses.create({
model: "grok-4.7",
input,
tools,
stream: true,
});
const response = await stream
.on("text", (text) => process.stdout.write(text))
.on("client_tool_call", (call) => {
if (call.type === "function_call") toolRuns.push(runTool(call));
})
.done();
if (toolRuns.length === 0) break;
const toolOutputs = await Promise.all(toolRuns);
input.push(...response.toInput(), ...toolOutputs);
}
The handlers map is the list of functions the model may call. runTool() returns errors to the model instead of throwing, so the model can recover, and a failing tool can't crash the loop while the stream is still running.
With the shell tool, the model writes shell commands and your application runs them. Use it for agents that work on your machine, such as exploring a repository, running tests, or checking disk space. Each call arrives through the client_tool_call listener as a shell_call, with the commands in action.commands. The model writes these commands, so run them in a sandbox or container, or check each one before running it:
import { exec } from "node:child_process";
import { SpaceXAI } from "@xai-official/sdk";
const client = new SpaceXAI();
const stream = await client.responses.create({
model: "grok-4.7",
input: "How much free disk space does this machine have?",
tools: [{ type: "shell", environment: { type: "local" } }],
stream: true,
});
await stream
.on("client_tool_call", (call) => {
if (call.type === "shell_call") {
for (const command of call.action.commands) {
exec(command, (error, stdout, stderr) => console.log(stdout || stderr));
}
}
})
.done();
To let the model use the results, send them back in the next request, like the function outputs in the tool call loop: { type: "shell_call_output", call_id: call.call_id, output: [{ stdout, stderr, outcome: { type: "exit", exit_code: 0 } }] }.
To give the model skills, list them in the tool's environment. A skill is a directory with a SKILL.md file of instructions. The model sees each skill's name and description, and when a task matches one, it reads the SKILL.md through your shell tool and follows it. If you only allow certain commands, allow reading the skill's directory:
const shell: Tool = {
type: "shell",
environment: {
type: "local",
skills: [
{
name: "release-notes",
description: "Write release notes from this repo's git log in our house style.",
path: "./skills/release-notes",
},
],
},
};
Asked to write release notes, the model reads ./skills/release-notes/SKILL.md, runs the git log command it describes, and writes the notes in the format it specifies.
Besides your own functions, the API has built-in tools. Create them with the helpers from @xai-official/sdk/tools, which check each tool's options as you type. SpaceXAI runs these tools and includes their results in the response:
webSearch()(web_search) searches the web. Options includeallowed_domains,excluded_domains,user_location, andsearch_context_size.xSearch()(x_search) searches posts on X. Options includeallowed_x_handles,excluded_x_handles,from_date, andto_date.codeExecution()(code_interpreter) writes and runs Python code to answer the prompt.collectionsSearch()(file_search) searches thecollections listed invector_store_ids.imageGeneration()(image_generation) creates or edits images.mcp()(mcp) calls tools on the remote MCP server atserver_url, identified byserver_label.toolSearch()(tool_search) loads the definitions of tools markeddefer_: truewhen the model needs them, instead of putting every definition in the prompt.
Two kinds of tools run in your application instead: function for your own functions, as shown in Tools, and shell, where the model writes shell commands for your application to run, as shown in Shell commands. When you stream, calls to both arrive through the client_tool_call listener.
The helpers return plain tool objects, so you can also write { type: "web_search" } yourself. Tool autocompletes the known types and accepts any other type, such as a tool released after this SDK version, but it doesn't check options the way the helpers do. See the SpaceXAI documentation for each tool's options.
Add the web search tool when a prompt needs current information:
import { SpaceXAI } from "@xai-official/sdk";
import { webSearch } from "@xai-official/sdk/tools";
const client = new SpaceXAI();
const response = await client.responses.create({
model: "grok-4.7",
input: "What are the latest developments in commercial spaceflight?",
tools: [webSearch()],
});
console.log(response.toText());
console.log(response.usage.num_server_side_tools_used);
Search posts on X, optionally limited to certain accounts and dates:
import { xSearch } from "@xai-official/sdk/tools";
const response = await client.responses.create({
model: "grok-4.7",
input: "What has SpaceXAI announced on X this month?",
tools: [xSearch({ allowed_x_handles: ["xai"], from_date: "2026-09-01" })],
});
allowed_x_handles and excluded_x_handles each take up to 20 handles and can't be used together. Set enable_image_understanding or enable_video_understanding to let the model look at media in posts.
When you stream, each finished search reaches the server_tool_call listener as a custom_tool_call. Its name is the search that ran, such as x_keyword_search, and input holds the search arguments as a JSON string:
await stream
.on("server_tool_call", (call) => {
if (call.type === "custom_tool_call") console.log(call.name, call.input);
})
.done();
Let the model write and run Python for calculations and data analysis:
import { codeExecution } from "@xai-official/sdk/tools";
const response = await client.responses.create({
model: "grok-4.7",
input: "What is the standard deviation of 12, 15, 19, 22, and 31?",
tools: [codeExecution()],
});
The code runs in a sandbox with common libraries installed. The tool takes no options.
Search documents you've added to collections:
import { collectionsSearch } from "@xai-official/sdk/tools";
const response = await client.responses.create({
model: "grok-4.7",
input: "What does our refund policy say about digital purchases?",
tools: [collectionsSearch({ vector_store_ids: ["your-collection-id"], max_num_results: 10 })],
});
Give the model the tools of a remote MCP server. SpaceXAI connects to the server and calls its tools during the response:
import { mcp } from "@xai-official/sdk/tools";
const response = await client.responses.create({
model: "grok-4.7",
input: "What is the modelcontextprotocol/typescript-sdk repository for?",
tools: [mcp({ server_url: "https://mcp.deepwiki.com/mcp", server_label: "deepwiki" })],
});
The server must use the Streaming HTTP or SSE transport. Limit the model to some of the server's tools with allowed_tools, and pass credentials with authorization or headers. require_approval and connector_id aren't supported yet. For a server with many tools, set defer_: true on it and add toolSearch(), so the model loads only the tool definitions it needs.
Add the image generation tool to let the model create or edit images as one step of a response. Each image arrives as an image_generation_call output item whose result holds base64 image data:
import { writeFile } from "node:fs/promises";
import { isImageGenerationCall, SpaceXAI } from "@xai-official/sdk";
import { imageGeneration } from "@xai-official/sdk/tools";
const client = new SpaceXAI();
const response = await client.responses.create({
model: "grok-4.7",
input: "Generate an image of a corgi surfing a big wave, in the style of a Japanese woodblock print.",
tools: [imageGeneration()],
});
console.log(response.toText());
const call = response.output.find(isImageGenerationCall);
if (call?.result) {
await writeFile("corgi.jpg", Buffer.from(call.result, "base64"));
}
Pass action: "generate" or action: "edit" to imageGeneration() to allow only one of those capabilities. When streaming, each call emits response.image_generation_call.in_progress, response.image_generation_call.generating, and response.image_generation_call.completed events, then a response.output_item.done event carries the finished item.
To generate or edit an image directly with full control over its size and format, use the image generation and image editing methods instead.
Every completed response provides:
response.toText()to concatenate output textresponse.toInput()to carry all output items into a later requestresponse.toJson()to parse completed JSON outputresponse.parsedfor non-throwing JSON parsingresponse.outputfor typed output itemsresponse.usagefor token counts, server-side tool use, and cost when availableresponse.httpfor the HTTP status, headers, SpaceXAI request ID, and client request IDresponse.rawfor the response object as the API sent it, including fields this SDK doesn't know yet
Use the exported type guards when inspecting output items:
import {
isFunctionCall,
isMessage,
isReasoning,
} from "@xai-official/sdk";
for (const item of response.output) {
if (isMessage(item)) {
console.log("message", item.content);
} else if (isFunctionCall(item)) {
console.log("function", item.name, item.arguments);
} else if (isReasoning(item)) {
console.log("reasoning item", item.id);
}
}
Generate images from a text prompt with a Grok Imagine model. Images are returned as temporary URLs by default, so download or process them promptly:
const result = await client.images.generate({
model: "grok-imagine-image-2.0",
prompt: "A collage of London landmarks in a stenciled street-art style",
});
console.log(result.data[0]?.url);
console.log(result.usage?.cost_usd);
Request up to 10 images with n, and shape the output with aspect_ratio, resolution, and quality. Only grok-imagine-image-2.0 supports quality. Set response_format: "b64_json" to receive base64 data instead of URLs:
import { writeFile } from "node:fs/promises";
const result = await client.images.generate({
model: "grok-imagine-image-2.0",
prompt: "A futuristic city skyline at night",
n: 4,
aspect_ratio: "16:9",
resolution: "2k",
response_format: "b64_json",
});
for (const [index, image] of result.data.entries()) {
if (image.b64_json) {
await writeFile(`skyline-${index}.jpg`, Buffer.from(image.b64_json, "base64"));
}
}
Base64 output is about a third larger than the image file, so large batches of high-resolution images can approach the default 32 MiB response size limit. Raise maxResponseBodyBytes for those requests:
const result = await client.images.generate(
{
model: "grok-imagine-image-2.0",
prompt: "A futuristic city skyline at night",
n: 10,
resolution: "2k",
response_format: "b64_json",
},
{ maxResponseBodyBytes: 128 * 1024 * 1024 },
);
Each result provides data, usage, and http. usage.cost_usd converts the reported cost_in_usd_ticks to US dollars, and usage is null when the API omits it.
Pass a source image with your prompt to edit it. image accepts a public URL, a base64 data URL, a Files API file_id, or a Blob or File, which the SDK converts to a data URL before sending the request:
import { openAsBlob } from "node:fs";
const photo = await openAsBlob("./photo.png");
const result = await client.images.edit({
model: "grok-imagine-image-2.0",
prompt: "Render this as a pencil sketch with detailed shading",
image: photo,
});
console.log(result.data[0]?.url);
When a Blob or File has no MIME type, as with openAsBlob() or new File([bytes], "photo.png"), the SDK detects JPEG, PNG, or WebP from its first bytes. The API rejects image data URLs that are not typed as one of those formats.
To combine up to five source images, pass images instead of image and refer to them in the prompt as <IMAGE_0>, <IMAGE_1>, and so on. The output follows the first image's aspect ratio unless you set aspect_ratio:
const result = await client.images.edit({
model: "grok-imagine-image-2.0",
prompt: "Place the cat from <IMAGE_0> on the sofa from <IMAGE_1>",
images: [
{ url: "https://example.com/cat.png" },
{ file_id: "file_abc123" },
],
aspect_ratio: "16:9",
});
Video generation runs as a background job. generate() starts the job and returns its request_id, and wait() polls until the job finishes:
const { request_id } = await client.videos.generate({
model: "grok-imagine-video-1.5",
prompt: "A paper boat drifting down a rain-soaked street",
duration: 8,
aspect_ratio: "16:9",
resolution: "720p",
});
const result = await client.videos.wait(request_id);
if (result.status === "done") {
console.log(result.video?.url);
console.log(result.usage?.cost_usd);
} else {
console.error(result.status, result.error?.code, result.error?.message);
}
wait() resolves once the status is no longer pending: done, failed, or expired. A failed result includes an error with a code and message. If video.respect_moderation is false, the video did not pass moderation and has no URL. Video URLs are temporary, so download the file promptly.
wait() polls every 5 seconds for up to 10 minutes. Pass interval and timeout in milliseconds to change this, and a signal to stop waiting. A timeout rejects with TimeoutError but does not cancel the job, so you can call wait() again. To check once without waiting, call client.videos.get(request_id), which returns status: "pending" until the video is ready.
To animate a still image, pass it as image. image, reference_images, and keyframe images accept a public URL, a base64 data URL, a Files API file_id, or a Blob or File, which the SDK converts to a data URL before sending the request:
import { openAsBlob } from "node:fs";
const { request_id } = await client.videos.generate({
model: "grok-imagine-video-1.5",
prompt: "Make the water crash down and slowly pan out the camera",
image: await openAsBlob("./waterfall.png"),
});
Edit a video with edit(), or continue it from its last frame with extend(). Both return a request_id for wait(). The source video must be an MP4, given as a public URL, a base64 data URL, a Files API file_id, or a Blob or File, which the SDK converts to a data URL before sending the request. For extensions, duration sets the length of the new segment only:
import { openAsBlob } from "node:fs";
const edit = await client.videos.edit({
model: "grok-imagine-video",
prompt: "Give the woman a silver necklace",
video: await openAsBlob("portrait.mp4"),
});
const extension = await client.videos.extend({
model: "grok-imagine-video",
prompt: "The camera slowly zooms out to reveal the city skyline",
video: { file_id: "file_abc123" },
duration: 6,
});
A Blob or File without a MIME type is sent as video/mp4. Because the video travels inside the request as base64, which is a third larger than the file, upload large videos with client.files.upload() (up to 50 MB) and pass { file_id } instead.
List the video generation models available to your API key with client.models.video.list(), or look one up by ID with client.models.video.get().
Upload a document, image, or video once and refer to it by ID. A file ID works wherever the API accepts a file_id, such as an input_file part in the Responses API or an image or video input:
import { openAsBlob } from "node:fs";
const file = await client.files.upload({
file: await openAsBlob("./report.pdf"),
filename: "report.pdf",
});
const response = await client.responses.create({
model: "grok-4.7",
input: [
{
role: "user",
content: [
{ type: "input_text", text: "Summarize the key findings in this report." },
{ type: "input_file", file_id: file.id },
],
},
],
});
console.log(response.toText());
The API records the upload's filename as the file's filename. A File uses its own name, and a plain Blob, such as one from openAsBlob(), needs filename. Files are kept until you delete them; set expires_after to between 3,600 and 2,592,000 seconds (1 hour to 30 days) to have one deleted automatically.
List, download, share, and delete stored files:
import { writeFile } from "node:fs/promises";
for await (const stored of client.files.list()) {
console.log(stored.id, stored.filename, stored.bytes);
}
const content = await client.files.content(file.id);
await writeFile("report-copy.pdf", await content.bytes());
const { public_url } = await client.files.createPublicUrl(file.id, {
expires_after: 86_400,
});
console.log(public_url);
await client.files.revokePublicUrl(file.id);
await client.files.delete(file.id);
list() returns the newest files first and fetches further pages as the loop needs them. content() returns an BinaryResponse: stream its body or read it with bytes(), text(), or blob().
Anyone with a public URL can download the file without an API key. Only images, videos, and PDFs up to 50 MiB can be made public. A file has at most one public URL, so calling createPublicUrl() again returns the existing URL and updates its expiry if you pass a new expires_after. Without expires_after, the URL lasts as long as the file unless you revoke it. After revoking, copies already cached by the CDN can still be served briefly.
The Batch API processes large volumes of requests asynchronously at a reduced price. Most requests complete within 24 hours. Create a batch, then add requests to it:
const batch = await client.batches.create({ name: "feedback_sentiment" });
const feedback = [
{ id: "feedback_001", text: "The product exceeded my expectations!" },
{ id: "feedback_002", text: "Shipping took way too long." },
];
await client.batches.requests.add(batch.batch_id, {
batch_requests: feedback.map((item) => ({
batch_request_id: item.id,
batch_request: {
responses: {
model: "grok-4.7",
input: [
{ role: "system", content: "Classify the sentiment as positive, negative, or neutral." },
{ role: "user", content: item.text },
],
},
},
})),
});
Each batch_request holds one request. responses takes the same CreateParams as client.responses.create(), including the store: false default, and its result comes back as a chat_get_completion response. image_generation, image_edit, video_generation, and video_extension take the request body of the matching REST endpoint. Results can come back in any order, so give each request a batch_request_id that is unique within the batch. Not every model accepts batch requests; each model page lists its Batch API support.
Wait until no requests are pending, then read the results:
await client.batches.wait(batch.batch_id);
for await (const { batch_request_id, batch_result } of client.batches.results(batch.batch_id)) {
if ("error" in batch_result) {
console.error(batch_request_id, batch_result.error);
} else {
console.log(batch_request_id, batch_result.response);
}
}
wait() polls every 5 seconds and rejects with TimeoutError after 24 hours. Pass interval, timeout, or signal to change that. Results are available as soon as each request finishes, so you can read them before the whole batch completes. Use client.batches.requests.list() to check the state of individual requests, client.batches.list() to list your team's batches, and client.batches.cancel() to stop the remaining requests. Finished results stay available after cancelling.
Convert text to speech with client.voice.speak(). The audio comes back as an BinaryResponse, encoded as MP3 unless you set output_format:
import { writeFile } from "node:fs/promises";
const speech = await client.voice.speak({
text: "Welcome to SpaceX. [] How can I help you today?",
language: "en",
voice_id: "eve",
});
await writeFile("welcome.mp3", await speech.bytes());
Shape the delivery with speech tags in the text. Inline tags such as [], [long-], and [laugh] go where the sound should happen, and wrapping tags such as <whisper>It's a secret.</whisper> change how the enclosed text is spoken. The API doesn't report mistakes in tags, so TypeScript checks string literals as you type: it flags unknown tags such as [laff] and suggests the closest one, and it catches wrapping tags that are never closed, closed without being opened, or closed in the wrong order. To use a tag released after this SDK version, add as UnsafeSpeechText to the text, which skips the check. Searching for UnsafeSpeechText then finds every tag to clean up once the SDK knows it:
import { type UnsafeSpeechText } from "@xai-official/sdk";
await client.voice.speak({
text: "Hello [new-tag] there." as UnsafeSpeechText,
language: "en",
});
voice_id autocompletes the built-in voices and accepts any other string, such as a custom voice ID or a voice added after this SDK version. List the built-in voices with client.voice.list(). To start playback before synthesis finishes, read speech.body as a stream. Set with_timestamps: true to receive JSON with base64 audio and per-character audio_timestamps instead of audio bytes.
Transcribe a recording with client.voice.transcribe(). Pass the audio as a Blob or File, or pass url to have the API download it:
import { openAsBlob } from "node:fs";
const transcript = await client.voice.transcribe({
file: await openAsBlob("./meeting.mp3"),
language: "en",
format: true,
});
console.log(transcript.text);
format: true writes spoken numbers, currencies, and units in written form, and requires language. Word-level timings are in transcript.words.
Clone a voice from a reference clip of up to 120 seconds with client.voice.custom.create(). Creating custom voices through the API requires an Enterprise plan:
import { openAsBlob } from "node:fs";
const voice = await client.voice.custom.create({
file: await openAsBlob("./reference.wav"),
name: "Friendly Narrator",
language: "en",
});
console.log(voice.voice_id);
Pass the returned voice_id to speak() or a realtime session like a built-in voice. client.voice.custom also provides list(), get(), update(), delete(), and getAudio(), which downloads the reference clip.
Realtime voice sessions in a browser should authenticate with a short-lived client secret instead of your API key. Create one on your server:
const secret = await client.voice.clientSecrets.create({
expires_after: { seconds: 300 },
});
Send secret.value to the browser, which passes xai-client-secret.<value> as the WebSocket subprotocol when it connects to wss://api.x.ai/v1/realtime. Secrets expire after 10 minutes by default, and expires_after.seconds can be at most 3600.
Encode text with a language model's tokenizer to count its tokens or see how it is split:
const { token_ids } = await client.tokenizer.encode({
model: "grok-4.7",
text: "Hello world!",
});
console.log(token_ids.length);
for (const token of token_ids) {
console.log(token.token_id, token.string_token);
}
Inference requests add tokens of their own, so usage.input_tokens for a prompt can be higher than this count.
The API has no decode endpoint, but each token carries its bytes, so you can turn encoded tokens back into text. Decode token_bytes rather than joining string_token, because a token can hold part of a multi-byte character:
const text = new TextDecoder().decode(
new Uint8Array(token_ids.flatMap((token) => token.token_bytes)),
);
Use model IDs directly. ModelId suggests known string literals while still accepting models released after the installed SDK version:
import { type KnownModelId } from "@xai-official/sdk";
const model: KnownModelId = "grok-4.7";
const available = await client.models.list();
for (const availableModel of available.data) {
console.log(availableModel.id);
}
const modelInfo = await client.models.get(model);
console.log(modelInfo);
KnownModelId is generated from the SpaceXAI model documentation. Use it when you want strict validation against the models known to this SDK release.
Image generation models have their own catalog, which includes modalities, aliases, and pricing. ImageModelId and KnownImageModelId work the same way as the text model types:
import { type KnownImageModelId } from "@xai-official/sdk";
const imageModel: KnownImageModelId = "grok-imagine-image-2.0";
const imageModels = await client.models.image.list();
for (const availableImageModel of imageModels.models) {
console.log(availableImageModel.id, availableImageModel.aliases);
}
const imageModelInfo = await client.models.image.get(imageModel);
console.log(imageModelInfo);
Chat and image understanding models have their own catalog, which includes modalities, aliases, token pricing, and supported reasoning efforts:
const languageModels = await client.models.language.list();
for (const languageModel of languageModels.models) {
console.log(languageModel.id, languageModel.input_modalities);
}
const languageModelInfo = await client.models.language.get("grok-4.7");
console.log(languageModelInfo.capabilities?.reasoning_effort);
Token prices such as prompt_text_token_price are in USD cents per 100 million tokens. Divide them by 10,000 for US dollars per million tokens.
client.account.apiKey() returns the name, status, and permissions of the API key the client is using:
const apiKeyInfo = await client.account.apiKey();
console.log(apiKeyInfo.name, apiKeyInfo.acls, apiKeyInfo.api_key_disabled);
The SDK sends store: false unless you opt in. This differs from the API wire default. With storage disabled, the SDK requests encrypted reasoning content so response.toInput() can preserve context between turns.
Pass store: true when you need to retrieve, continue, inspect, or delete a response by ID:
const stored = await client.responses.create({
model: "grok-4.7",
input: "Save this response.",
store: true,
});
const fetched = await client.responses.get(stored.id);
const inputItems = await client.responses.inputItems.list(stored.id);
await client.responses.delete(stored.id);
List methods that return results in pages fetch the next page for you in a for await loop:
for await (const file of client.files.list()) {
console.log(file.id);
}
This works for files.list(), batches.list(), batches.results(), batches.requests.list(), voice.custom.list(), and responses.inputItems.list(). Awaiting one of these calls instead returns a single page.
Configure defaults on the client:
const client = new SpaceXAI({
timeout: 60_000,
idleTimeout: 30_000,
maxRetries: 2,
});
Override them for one request and pass an AbortSignal when needed:
const controller = new AbortController();
const pending = client.responses.create(
{
model: "grok-4.7",
input: "Write a detailed report.",
},
{
signal: controller.signal,
timeout: 120_000,
maxRetries: 0,
},
);
controller.abort();
await pending;
Requests that generate content, such as responses.create, images.generate, and images.edit, retry only explicit 429 responses by default. Read-only requests may also retry transient HTTP failures. Retry delays honor Retry-After and use jittered exponential backoff.
When you leave out stream, responses.create() streams the response under the hood and resolves to the final response. Streamed responses send headers right away, so long reasoning requests aren't cut off by limits on waiting for headers, such as the 5 minutes that Node's built-in fetch allows whatever timeout is set to. Reasoning can also go quiet for minutes, so these requests only apply idleTimeout when you pass it on the request. Set stream: false to send a plain JSON request instead.
All SDK errors extend APIError. Status-specific classes are exported for common API failures.
import {
APIError,
AuthenticationError,
RateLimitError,
} from "@xai-official/sdk";
try {
await client.responses.create({
model: "grok-4.7",
input: "Hello",
});
} catch (error) {
if (error instanceof AuthenticationError) {
console.error("Check XAI_API_KEY");
} else if (error instanceof RateLimitError) {
console.error("Rate limited. Retry later.");
} else if (APIError.is(error)) {
console.error(error.status, error.code, error.param, error.requestId, error.message);
} else {
throw error;
}
}
The SDK exports APIConnectionError, APIProtocolError, APIStatusError, AbortError, AuthenticationError, NotFoundError, OverloadedError, PermissionDeniedError, RateLimitError, and TimeoutError.
Set XAI_DEBUG=1 to print each request's method, URL, and headers as a cURL command:
XAI_DEBUG=1 node app.js
Authentication headers and common credential fields are redacted. Request bodies are always omitted because prompts and tool outputs may contain sensitive data.
Structured API failures expose error.type, error.code, and error.param when the server returns them. The SpaceXAI request ID is also available at response.http.requestId and error.requestId. Include it when reporting an API problem.
Every request also sends an x-client-request-id header with a UUID generated by the SDK. The ID stays the same across retries and is available at response.http.clientRequestId and error.clientRequestId, even when a request fails before the API responds. To use your own ID, set x-client-request-id in the request headers.
Install dependencies and run the checks:
pnpm install
pnpm check
pnpm pack:check
pnpm pack:smoke
Run every release gate, including the dependency audit:
pnpm release:check
Create a distributable tarball and SHA-256 checksum in artifacts/:
pnpm pack:artifact
Generated API types live in src/generated/types.ts, speech tags, voice IDs, and voice model IDs in src/generated/voice.ts, and the model ID union in src/models.ts. Run pnpm generate:types for API types, pnpm generate:voice for the Voice API values, and pnpm generate:models for model IDs instead of editing those files by hand.
We aren't accepting outside contributions yet, but plan to later. Please open an issue for bugs and feature requests. CONTRIBUTING.md covers local development.
Report suspected vulnerabilities privately as described in SECURITY.md. Do not include credentials, confidential data, or unredacted request bodies in public issues.
Licensed under the Apache License 2.0.