Gemini 3.8 text-to-speech says hello Google introduced two new text-to-speech models in its Gemini family, Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, expanding voice generation from 30 original voices to a library of 2,000+ production-ready voices across more than 100 languages and dialects. Gemini 3.8 Flash TTS is built for creative direction and character design with generative voice design and voice replication from a 30-second audio sample, while Gemini 3.8 Flash-Lite TTS targets high-volume, cost-efficient dubbing and voice agents. Both models offer line-by-line performance control, long-form generation with minimal speaker drift, and native two-speaker scene staging, with voice replication backed by consent verification, SynthID watermarking, and C2PA credentials. Gemini 3.8 text-to-speech says hello Today, we’re introducing two new text-to-speech models to the Gemini family, transforming voice generation from static presets into a dynamic creative studio. These models enable creators, developers, and enterprises to create richer, more expressive audio experiences, while enabling improved user experiences in products like Gemini Notebook https://notebook.google.com/ and Google Vids http://vids.new/ . - Gemini 3.8 Flash TTS: Built for deep creative direction and character design. Create entirely new voices from scratch using natural language prompts to bring characters to life across gaming, immersive audiobooks, podcasts, and interactive media. Direct every performance line by line with granular control over acting cues, pacing, dialect shifts, and backchanneling. - Gemini 3.8 Flash-Lite TTS: Built for high-volume, cost-efficient scale. Optimized for high-volume dubbing, audio content creation, and expressive voice agents with fine-grained control over tone, pacing, and expressive nuance. These models complement our fast-growing Gemini Audio family, following 3.5 Live Translate https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-live-3-5-translate/ , 3.5 Transcribe https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5-transcribe/ , 3.8 Live, and 3.8 Live Extended Thinking https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-live-gemini-3-8-live-extended-thinking/ . Create and customize your own voices Scale up from 30 original voices to an infinite library. Whether you need an entirely original character voice or a consistent brand ambassador, our 3.8 Flash TTS model powers a full vocal studio. This enables you to create and use expressive, natural-sounding voices for every moment, while empowering developers and enterprises to easily build custom audio experiences. - Generative voice design: With Gemini 3.8 Flash TTS, create bespoke voices from scratch by customizing role, accent and voice characteristics across more than 100 languages and dialects using natural language prompting — whether you're bringing a dramatic, fire-breathing dragon to life or crafting a charismatic narrator with a distinct regional cadence. - Expansive voice library: Access 2,000+ production-ready voices with broad language coverage — including regional varieties like Mexican Spanish, Quebec French, and Scots English. - Voice replication: Recreate consistent vocal profiles from just a 30-second audio sample of your voice or a voice you have the rights to use, backed by built-in consent verification, SynthID watermarking, and C2PA credentials to protect both developers and their vocal talent. - Save and scale: Save and manage the custom voices you designed to ensure consistent performance and minimal drift across ongoing projects. - Voice remixing: Coming soon, pick a voice from our voice library and fine-tune timbre, pitch, pace, and accent. Use prompts to dial in characteristics e.g. “add subtle Southern US accent” or “soften the delivery” . Direct the performance, line by line Once you've selected your voices, both TTS models give you precise control over how each line is delivered. - Direct performance line by line: Write your own stage directions or let Gemini steer delivery with natural script cues — from a calm customer service agent to a whispered suspense scene. - Long-form generation: Maintain high voice quality, natural pacing, and character timbre across hours of continuous audio with minimal speaker drift — ideal for podcasts and audiobooks. - Native two-speaker scene staging: Direct multi-turn conversations seamlessly from a single script —whether for a podcast or dramatic storytelling—while keeping both voices distinctly separated with natural conversational turn-taking. - Scripted vocal bursts & backchanneling: Add realistic conversational texture using non verbal cues like