Build realtime voice agents on AI Gateway

wpnews.pro

cd /news/artificial-intelligence/build-realtime-voice-agents-on-ai-ga… · home › topics › artificial-intelligence › article

[ARTICLE · art-43621] src=vercel.com ↗ pub=2026-06-29T07:00Z topic=artificial-intelligence verified=true sentiment=↑ positive

Build realtime voice agents on AI Gateway

Vercel's AI Gateway now supports realtime voice, text-to-speech, and speech-to-text, enabling developers to build voice agents with the same API used for other modalities. The feature, in beta with AI SDK 7, offers models from OpenAI and xAI, with provider routing, observability, and spend controls. This allows low-latency, two-way voice conversations, voiceovers, and transcription, differentiating from pipeline approaches by using a single realtime model that processes audio directly.

read3 min views1 publishedJun 29, 2026

AI Gateway now supports audio/voice. You can add realtime voice, text to speech, and speech to text with the same calls you already use for text, image, and video, routed through AI Gateway alongside every other modality.

Audio launches with models from OpenAI and xAI. Each call gets the same provider routing, observability, spend controls, and bring-your-own-key support you already use for your other models.

These capabilities are in beta and available in AI SDK 7. | | | |---|---|---| | Live audio in and out, for streaming, low-latency session | Two-way voice agents and live conversation | | Text in, audio file out, single request | Voiceovers, spoken responses, audio versions of written content | | Recorded audio in, text out, single request | Transcribing voice notes, call recordings |

Realtime, speech, and transcription model are supported on AI SDK 7. Realtime turns your app into something a user can hold a conversation with. When they speak, the model responds right away. Because it replies in the moment instead of waiting for a full turn, users can interrupt and talk over it the way they would with a person. It fits voice assistants, customer support agents, hands-free tools, and anywhere a user would rather talk than type.

What sets it apart from chaining models together is that a single realtime model hears audio and produces audio directly, instead of running a speech-to-text, then language model, then text-to-speech pipeline.

In the browser, the useRealtime

hook manages the WebSocket connection, microphone capture, and audio playback.

The connection is authenticated with your AI Gateway credential, so you mint a short-lived token on the server and hand the browser only that token. Your API key never reaches the client. Add a route that mints the token:

Then connect from a client component:

The hook captures the microphone, streams the audio to the model through AI Gateway, and plays back the spoken reply. Outside the browser, you can drive the session over a WebSocket yourself with getWebSocketConfig

, serializeClientEvent

, and parseServerEvent

. See the realtime reference for that path. A realtime session works differently from a normal model call:

**Turn-taking and interruptions. **turnDetection: { type: 'server-vad' } lets the server decide when the user has stopped speaking, and lets the user talk over the model to cut a reply short (barge-in), with no client-side silence timers.

Tools mid-conversation. The model emits a tool call mid-reply, you run it and return the result as a client event, and it folds the answer into what it says next without ending the turn.

Generate spoken audio from text with generateSpeech

. Pass a voice and an output format, then write the result to a file:

Transcribe recordings into text with transcribe

. The audio can be a buffer, a base64 string, or a URL:

Speech and transcription are complementary, so they compose. You can generate audio with one model and read it back with the other, which is a quick way to check both ends of an audio pipeline.

You can also try audio models without writing any code. Open the models page, click into a model, and interact with it right in the browser. Talk to a realtime model to hold a voice conversation, or send text or audio to a speech or transcription model and read or play back the result.

Audio calls behave like every other model call on AI Gateway. You use one API key across providers, see requests and usage in observability, apply the same budgets and spend limits, and bring your own provider keys when you need to. Adding speech to an app that already uses AI Gateway for text/images/videos can now all be done in the same place.

source & further reading

vercel.com — original article Query Speed Insights from the Vercel CLI Realtime voice, speech, and transcription now supported on AI Gateway Add MCP Apps to Your AI SDK Application

~/api · this article 200

$curl api.wpnews.pro/v1/news/build-realtime-voice-age…

Read original on vercel.com → vercel.com/blog/realtime-voice-agents-on-ai-gate…

mentioned entities

Vercel

AI Gateway

OpenAI

xAI

AI SDK 7

metadata

slugbuild-realtime-voice-agents-on-ai-gateway

topic#artificial-intelligence

secondary4 topics

sentimentpositive

canonicalvercel.com

navigation

← prevPeeter P. Mõtsküla: AI agents ne…

next →Spec-Driven Development vs Vibe …

── more in #artificial-intelligence 4 stories · sorted by recency

devclubhouse.com · 29 Jun · #artificial-intelligence

The Death of the Single-Model API Call

dev.to · 29 Jun · #artificial-intelligence

Building a Legal AI Platform on Aurora DSQL and Vercel

restofworld.org · 29 Jun · #artificial-intelligence

Low-cost Chinese AI models like DeepSeek gain traction in the U.S.

marcoziccardi.com · 29 Jun · #artificial-intelligence

An agent opened this pull request. Nobody asked it to

── more on @vercel 3 stories trending now

wpnews · 28 May · #ai-startups

[AINews] Cognition raises $1B in $26B Series D

wpnews · 5 Jun · #ai-agents

Miasma Worm Targets AI Coding Agents via GitHub Repos

wpnews · 28 May · #ai-startups

The Niche SaaS Opportunity Map 2026: Highly Demanded Subscribed Categories Beyond Mainstream

sponsored brought to you by zahid.host 4,200+ EU-deployed projects

reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main

→ Live at https://your-agent.zahid.host ✓

Get free account → Pricing

from €0/mo · no card required