Google today introduced Gemini 3.5 Transcribe as its “most precise speech-to-text model yet” that is already powering several first-party products.
Unlike conventional speech recognition models that struggle with background noise, complex jargon, and disfluency cleanup, Gemini 3.5 Transcribe converts raw audio directly into accurate, polished, formatted text.
This model is “designed to capture your natural speaking style to better understand your intent and recognize custom vocabulary.” As seen in Rambler, Gemini 3.5 Transcribe can handle self-corrections (*“*let’s meet Tuesday—no, Wednesday”) and remove “ums,” “ahs,” and other filler words from the end result, as well as auto-format your text and allow for natural voice editing. Additionally, Google touts:
More precise transcription: As measured by Artificial Analysis, achieves an average Word Error Rate (WER) of 4.0% for streaming and 2.6% for non-streaming use-cases. It shows strong performance across noisy, real-world environments, accurately capturing alphanumeric entities like postal codes and order IDs.Custom vocabulary: Recognizes specialized jargon and unique spellings by seamlessly adapting transcriptions to your provided custom vocabulary.Global language support: Automatically detects and transcribes over 85 languages, seamlessly handling regional accents and diverse dialects.Multi-speaker identification: Accurately attributes speech in pre-recorded audio with timestamps for up to three speakers (support for 3+ speakers is experimental).
Performance-wise, Gemini 3.5 Transcribe offers what Google calls a “major advancement” across capabilities, improved word error rates, and significantly better latency over its Chirp 3 transcription model from 2025.
As measured by Artificial Analysis, time to final transcription, for example, improves by 70%. On the FLEURS benchmark across a set of top languages and locales, the model delivers precise multilingual performance, improving over Chirp 3, and achieving a 5.50% WER in streaming mode and 5.04% WER in non-streaming use-cases.
The other goal is to let you “execute tasks with your voice.” Function calling allows Gemini 3.5 Transcribe to “delegate complex tasks (such as image generation and file analysis) to other Gemini models.” This can be seen with the Speak to Window capability in the Gemini app for macOS.
Besides the Gemini macOS app and Gboard Rambler on Android, Gemini 3.5 Transcribe is available in Google Antigravity’s prompt box microphone where it “pairs screen context and chat history, with your permission, to ensure pinpoint transcription accuracy across file names, agent thoughts, and active documents.”
It’s coming next to the Chrome browser so you can “talk to type in any web field — making it effortless to dictate replies, draft posts, or prompt Gemini in Chrome more naturally and easily with your voice.”
Gemini 3.5 Transcribe is available:
For developers: In public preview in theGemini API via Google AI StudioandGoogle Antigravity.For enterprises: In public preview viaGemini Enterprise Agent Platformand coming soon toGemini Enterprise for Customer Experience.
*FTC: We use income earning auto affiliate links.* [More.](https://9to5mac.com/about/#affiliate)
[our homepage](http://9to5google.com/)for all the latest news, and follow 9to5Google on
[exclusive stories](https://9to5google.com/feature/exclusive/),
[reviews](https://9to5google.com/guides/reviews/),
[how-tos](https://9to5google.com/guides/how-to/), and
[subscribe to our YouTube channel](https://www.youtube.com/9to5google)