Google Releases Gemini 3.8 Live-Extended Conversational Model, Claims Better Performance Than Rivals At Lower Price Google DeepMind released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two live audio models that Google says top Artificial Analysis' Speech to Speech Quality Index at 82.6% and cost $0.84 and $3.50 per hour of input audio respectively, versus $5.83 an hour for GPT-Live-1 Astra and $4.80 for Grok Voice Think Fast 2.0. Both models switch between 97 languages mid-conversation, run tool calls and API requests in the background, and process visual input in near real time, with rollout starting today via the Gemini API, Google AI Studio, and Search Live, plus Gemini Live and Google Workspace apps for AI Pro and Ultra subscribers. Google also named LiveKit, Pipecat, Agora, Salesforce, and ServiceNow as partners building voice agents on the models, and said all audio output carries its SynthID watermark. Google DeepMind has rolled out two new live audio models, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, its latest attempt to make voice agents feel less like scripted IVR menus and more like an actual conversation partner. The launch follows Google’s now-familiar pattern of shipping model after model in quick succession this year — the company released Gemini 3.8 Flash https://officechai.com/ai/google-is-internally-testing-out-gemini-3-8-flash-reports/ just two weeks ago, and before that, Gemini 3.7 Flash https://officechai.com/ai/gemini-3-7-flash-benchmarks/ — but this time the focus has shifted to the audio side of the Gemini stack. Gemini 3.8 Live is pitched as the scale-and-cost-efficiency option, built for fluid dialogue and visual grounding, while 3.8 Live Extended Thinking is aimed at higher-complexity tasks that need more reasoning without breaking the flow of conversation. Both models can detect and switch between 97 languages mid-conversation, execute tool calls and API requests in the background while continuing to talk, and process visual input in near real time — Google’s demo videos show the model narrating a chess game and walking someone through an onboarding flow, using what’s on screen as context. On the benchmarks that matter most for voice agents, Google says 3.8 Live Extended Thinking takes the top spot on Artificial Analysis’ Speech to Speech Quality Index with 82.6%, ahead of GPT-Live-1 Astra Medium at 81.5% and Grok Voice Think Fast 2.0 High at 81.3%. The gap widens further on agentic performance: 3.8 Live Extended Thinking scores 68.6% on Artificial Analysis’ τ-Voice benchmark, narrowly ahead of GPT-Live-1 Astra’s 67.9% and comfortably ahead of Grok Voice Think Fast 2.0’s 56.5%. On Sierra’s τ³-Banking Leaderboard, which tests how well a voice agent can actually resolve customer-service style banking tasks, 3.8 Live Extended Thinking leads with 35.1%, versus 32.0% for GPT-Live-1 Astra and 16.5% for xAI-Realtime. The pricing story is where Google is leaning hardest. On Artificial Analysis’ cost-per-hour-of-input-audio measure using the Big Bench Audio subset, Gemini 3.8 Live comes in at just $0.84 an hour, and even 3.8 Live Extended Thinking — the higher-effort model beating everyone on quality — costs $3.50 an hour. That compares to $4.80 an hour for Grok Voice Think Fast 2.0 and $5.83 an hour for GPT-Live-1 Astra, meaning Google’s flagship reasoning model still undercuts its two closest rivals on cost while topping them on quality. The standard Gemini 3.8 Live model, meanwhile, is roughly a sixth of the price of GPT-Live-1 Astra for a good chunk of the same conversational ability, scoring 76.0% on the Speech to Speech Index. Google is rolling out 3.8 Live starting today through the Gemini API, Google AI Studio, and Search Live, with enterprise access coming via Gemini Enterprise. 3.8 Live Extended Thinking gets the same developer and enterprise rollout, plus a consumer release inside Gemini Live and Google Workspace apps — Docs, Gmail, and Keep — for Google AI Pro and Ultra subscribers. The company says it’s also partnering with platforms like LiveKit, Pipecat, and Agora, along with enterprise customers including Salesforce and ServiceNow, to build voice agents on top of the new models. All audio output from the models carries Google’s SynthID watermark.