Google unveils Gemini 3.8 Flash TTS with advanced voice customization Google launched its Gemini 3.8 Flash TTS models on September 23, offering developers 30 prebuilt studio voices, a custom voice design system driven by natural language prompts, and support for up to 130 languages through the Gemini API and Google AI Studio. The release includes two variants, gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts, and follows Google's September 15 launch of the Gemini 3.8 Live and 3.8 Live Extended Thinking models for real-time speech-to-speech interaction. The 3.8 generation arrives roughly five months after Gemini 3.1 Flash TTS in April 2026, with each release expanding the voice library and refining natural language control. Google 2015 logo Wikimedia Commons, public domain Google unveils Gemini 3.8 Flash TTS with advanced voice customization The new text-to-speech models offer 30 prebuilt voices, custom voice design, and support for 130 languages across podcasts, audiobooks, and interactive agents Google https://cryptobriefing.com/markets/alphabet/ just made it a lot harder to tell whether the voice reading your audiobook belongs to a human or a very well-trained language model. The company launched its Gemini 3.8 Flash TTS models on September 23, giving developers an expansive toolkit for generating synthetic speech that sounds less like a GPS navigator and more like an actual person with opinions about inflection. The release includes two model variants: gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts . Both are available through the Gemini API and Google AI Studio, positioning them squarely at developers building voice-driven applications at scale. What the new models actually do At the core of the 3.8 Flash TTS lineup is a layered approach to voice selection. Users get three tiers to work with: 30 prebuilt studio voices for quick deployment, an Extended Voice Library with hundreds of additional options, and a custom voice design system that lets you describe the voice you want using natural language prompts. That last feature is worth pausing on. Instead of tweaking sliders for pitch and speed like some 2019-era audio editor, developers can write something closer to a character description. The system then generates a voice to match and assigns it a persistent voice … ID, meaning you can summon the same synthetic personality across sessions without recreating it from scratch. Voice replication rounds out the customization suite, letting users clone specific vocal characteristics for consistent output. The models also support both single-speaker and multi-speaker audio generation, with speech metadata providing granular control over style, accent, and emotional tone. Multilingual support spans up to 130 languages, which is the kind of number that matters when you’re building products for markets beyond English-speaking countries. AI, tech, and the markets they move—in one daily briefing. Daily. Free. Join 34,000+ readers across crypto, finance, and policy. Live models for real-time conversation The TTS launch didn’t arrive alone. Eight days earlier, on September 15, Google released its Gemini 3.8 Live and 3.8 Live Extended Thinking models. These are designed for a fundamentally different use case: real-time speech-to-speech interaction. Where the TTS models prioritize high-fidelity audio generation with detailed controllability, the Live models are built for back-and-forth dialogue. They’re aimed at interactive voice agents, the kind of systems that need to listen, process, and respond within the cadence of a natural conversation rather than simply reading a script aloud. The Extended Thinking variant adds a reasoning layer to that real-time loop, letting voice agents pause and think before they speak. Both model families share the same API infrastructure and multilingual foundation, giving developers the option to mix pre-generated narration with live conversational capabilities in a single application stack. The progression from 2.5 to 3.8 Google has been on a steady cadence with its TTS releases. The 2.5 series received updates in December 2025, followed by Gemini 3.1 Flash TTS in April 2026. The jump to 3.8 represents roughly five months of iteration, with each generation expanding the voice library and refining the natural language control interface. Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy https://cryptobriefing.com/editorial-policy/ .