Google DeepMind has introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two audio-focused models designed for more natural, near real-time conversations with AI. The models expand the Gemini Audio family beyond transcription, translation, and text-to-speech capabilities, with one aimed at high-volume voice interactions and the other built for more demanding reasoning while a conversation continues.
For businesses, the development matters because voice can become a more practical interface for AI-supported work. Rather than requiring users to type every instruction, a voice system could support conversational task handling, customer-facing interactions, or guided internal workflows. The important distinction is that the two models serve different levels of complexity, and Google has not published explicit public pricing for either model on its product page. Google describes the models on its official Gemini Audio page, which also identifies access routes through Google AI Studio, the Gemini Live API, the Gemini API, the Gemini app, and related Google services.
Gemini 3.8 Live is positioned for near real-time voice interfaces. Google describes it as providing conversational capabilities optimized for high-volume, cost-effective use, alongside near-real-time reasoning. That positioning makes it the more direct fit for applications where responsiveness and the ability to handle many interactions are central requirements.
Gemini 3.8 Live Extended Thinking is the higher-end option. Google says it is intended for complex reasoning and can orchestrate multiple agents to solve background tasks while maintaining a natural conversation. In practical terms, this points to voice interactions that do more than answer a straightforward spoken request. A system could continue speaking naturally with a user while coordinating a more involved task in the background.
| Model | Google's stated focus | Relevant workflow implication |
|---|---|---|
| Gemini 3.8 Live | Near real-time voice interfaces, conversational capabilities, high-volume and cost-effective use | Potential fit for responsive voice interactions at scale |
| Gemini 3.8 Live Extended Thinking | Complex reasoning and orchestration of multiple agents for background tasks during conversation | Potential fit for voice-led workflows that require more involved task coordination |
The table reflects Google's stated positioning, not a benchmark comparison. The official material does not provide performance measurements, token costs, latency figures, or a public price list that would allow a more detailed operational comparison.
The new live models sit within a broader Gemini Audio portfolio. Google's product page also highlights Gemini 3.5 Transcribe, Gemini 3.5 Live Translate, and Gemini 3.1 Flash TTS. Those capabilities address distinct audio tasks: turning speech into text, translating live audio, and generating speech.
Gemini 3.8 Live and Extended Thinking represent a different emphasis. Their focus is the conversational layer itself, especially real-time dialogue and reasoning. This matters for teams assessing voice AI because a transcription or text-to-speech feature alone is not the same as a system designed to hold a live exchange and act on the context of that exchange.
Google also highlights SynthID watermarking across its audio work. SynthID is intended to flag whether speech has been AI-generated or edited. For organizations considering generated voice in customer or employee interactions, that safety capability is relevant, although the product page does not set out a complete implementation process or policy framework for individual use cases.
Google identifies multiple paths to the new capabilities. Developers can explore them through Google AI Studio and build through the Gemini Live API and Gemini API. Consumer access is also referenced through the Gemini app and related Google services. These routes indicate that Gemini Audio is intended to span experimentation, development, and end-user experiences rather than remain limited to a single product surface.
What remains less clear is the commercial and technical detail a business would need before committing to a production deployment. The official page does not list explicit public pricing for Gemini 3.8 Live or Gemini 3.8 Live Extended Thinking. It also does not provide the operational details needed to determine which model will be more economical for a particular workload.
Before building around a voice model, teams should establish:
The multi-agent capability associated with Extended Thinking is particularly notable, but it should not be read as a ready-made business process. The value will depend on how reliably an implementation connects the model to the specific tools, data, and steps that make up the underlying work.
Voice AI creates an opportunity to reduce friction in tasks that begin with a spoken request, but a useful deployment still requires a clear handoff between conversation and action. That could mean routing information to an existing application, triggering an approved workflow, or keeping a user involved where a decision needs review.
Voice interfaces are only valuable when they connect reliably to the work that follows the conversation. Scalevise helps businesses turn promising AI capabilities into practical processes, from mapping suitable voice-led tasks to connecting models with the tools teams already use. Our AI workflow automation service can help reduce manual handoffs and build workflows with clear operational value. Discuss an AI automation project with Scalevise.
What is Gemini 3.8 Live?
Gemini 3.8 Live is a Gemini Audio model for near real-time voice interfaces. Google describes it as offering conversational capabilities optimized for high-volume, cost-effective use and near-real-time reasoning.
What is Gemini 3.8 Live Extended Thinking?
Gemini 3.8 Live Extended Thinking is a Gemini Audio model aimed at more complex reasoning. Google says it can orchestrate multiple agents to solve background tasks while maintaining a natural conversation.
How can developers access Gemini 3.8 Live models?
Google identifies Google AI Studio, the Gemini Live API, and the Gemini API as developer access paths. The Gemini app and related Google services are also listed as consumer access routes.
Has Google published pricing for Gemini 3.8 Live and Extended Thinking?
The official Gemini Audio page does not include explicit public pricing for Gemini 3.8 Live or Gemini 3.8 Live Extended Thinking.
What safety feature does Google highlight for Gemini Audio?
Google highlights SynthID watermarking, which is designed to flag whether speech was AI-generated or edited.
Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking extend Google's audio strategy from individual speech tasks toward real-time, conversational AI. The division between high-volume live interaction and more complex reasoning gives developers a clearer starting point for evaluating voice-led workflows. Access is available through Google's developer and consumer channels, while pricing and deployment-specific performance details remain matters to confirm through the relevant Google portals and testing.