Quick summary - click for full details #
Concise summary of key highlights
In one line
Google launches Gemini 3.8 Live and 3.8 Live Thinking AI models for real-time audio reasoning and multi-step tasks.
Key points
• Gemini 3
Enables real-time audio transcriptions, reasoning, and translations for multi-utility apps with near-instant visual input processing.
• Gemini 3
Designed for high-complexity tasks with advanced multi-step reasoning and increased intelligence.
• Gemini 3
Dedicated speech-to-text model offering highly precise transcription across 85+ languages with low error rates.
• Voice command integration
Users can invoke Gemini AI via smartphones to perform complex multi-step tasks across Google and third-party apps using voice commands.
• Watermarking for AI content
All audio generated by the new models is watermarked with SynthID to detect AI-generated content and prevent misinformation.
Key statistics
4.0%
Average Word Error Rate (WER) for streaming transcription
2.6%
Average Word Error Rate (WER) for non-streaming transcription
85+
Number of languages supported by Gemini 3.5 Transcribe
97
Number of languages automatically detected mid-conversation
Processed with AI. Reviewed by DH Digital Team.
Published 17 September 2026, 05:28 IST