cd /news/artificial-intelligence/google-rolls-out-gemini-audio-to-imp… · home topics artificial-intelligence article
[ARTICLE · art-112075] src=androidauthority.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Google rolls out Gemini Audio to improve real-time dialogue and speech recognition

Google has introduced Gemini Audio, a trio of speech models including Gemini 3.5 Transcribe, Gemini 3.5 Live, and Gemini 3.5 Live Experimental, designed to improve real-time dialogue and speech recognition. Gemini 3.5 Transcribe, which replaces Chirp 3, achieves a 4% word-error-rate on streaming audio and 2.6% on pre-recorded files, supports over 85 languages, and is rolling out across Google products such as Search Live, Gemini Live, Docs, Keep, Gmail, the Gemini app, and Gboard.

read3 min views3 publishedAug 26, 2026
Google rolls out Gemini Audio to improve real-time dialogue and speech recognition
Image: Androidauthority (auto-discovered)

Affiliate links on Android Authority may earn us a commission. Learn more.

Aug 26, 2026 — 12:00 PM ET

  • Google has introduced Gemini Audio, a trio of new models built to improve dialogue and speech recognition.
  • The models include Gemini 3.5 Live, 3.5 Live Experimental, and 3.5 Transcribe.
  • The new family of AI models is rolling out for everyone in Search Live, Gemini Live, Docs, Keep, Gmail, the Gemini app, and Gboard.

It’s been a few months since Google debuted the first of its Gemini 3.5 models. With the rapid pace of AI development, the company has since released Gemini 3.6 and Gemini 3.7. However, the tech giant hasn’t completely moved on from the 3.5 family just yet, as it is introducing new members to the group today.

Google has announced a trio of new speech-related models: Gemini 3.5 Transcribe, Gemini 3.5 Live, and Gemini 3.5 Live Experimental. Together, these models make up what the company calls Gemini Audio. According to Google, these models were built to power real-time dialogue and speech recognition for more natural and responsive conversations.

Gemini 3.5 Transcribe: Advanced transcription #

Gemini 3.5 Transcribe is described as the Mountain View-based firm’s most precise speech-to-text model yet. It will replace the previous transcription model, Chirp 3. Google claims that Gemini 3.5 Transcribe offers better precision, deep context awareness, and automatic language detection across more than 85 languages.

Transcribe delivers a handful of key benefits, such as:

Smart transcription: Removes filler words (like “ums” and “ahs”), auto-formats your text, and edits naturally with just your voice.Function calling: The model can delegate complex tasks (such as image generation) to other Gemini models via function calls. While function calls are already available in the Gemini macOS app for developers, support is expected to come to the API soon.More precise transcriptions: Google claims its model delivers a 4% WER (word-error-rate) on streaming audio and 2.6% WER on pre-recorded files across diverse real-world conditions, including background noise and conversational AI interactions.Custom vocabulary: Can recognize specialized jargon and unique spellings and adapt transcriptions to your provided custom vocabulary.Global language support: Automatically detects and transcribes over 85 languages, can also handle regional accents and diverse dialects.Multi-speaker identification: Attributes speech in pre-recorded audio with word-level timestamps for up to three speakers (support for over three speakers is experimental).

We’ve already had a taste of this model through the new Rambler feature for Gboard, which is available in select countries and languages. Users will also find Gemini 3.5 Transcribe on Google Antigravity, and the Gemini app on macOS. Eventually, it will also head to Chrome, allowing you to talk to type in any web field.

Gemini 3.5 Live and Live Experimental #

As for Gemini 3.5 Live, this model is designed to handle mid-sentence interruptions and process live visuals. Additionally, it can blend multiple languages and trigger background tools. The model is capable of doing all of this without the need to the conversation.

Meanwhile, Live Experimental takes things a little further. This model is meant to handle more complex tasks, reasoning directly while speaking and narrating its progress step by step.

You’ll be able to start trying out the Gemini Audio family of models soon. Google says that these models are rolling out for everyone in Search Live, Gemini Live, Docs, Keep, Gmail, the Gemini app, and Gboard. Meanwhile, they’ll be available across the Gemini API via Google AI Studio and Google Antigravity for developers. And for enterprise customers, these models will come to the Gemini Enterprise Agent Platform and Gemini Enterprise for Customer Experience.

Thank you for being part of our community. Read our Comment Policy before posting.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @google 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/google-rolls-out-gem…] indexed:0 read:3min 2026-08-26 ·