Google launches Gemini 3.8 voice models with Hume founder Alan Cowen credited Google launched Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS on September 23rd, text-to-speech models that add custom voice design and line-by-line performance control, with former Hume AI founder and Google DeepMind Director of Research Science Alan Cowen credited alongside Leland Rechis. Google says the models support more than 100 languages and dialects, offer a library of more than 2,000 production-ready voices, and can replicate a voice from a 30-second audio sample, with replication unavailable in Illinois, Texas, the European Economic Area, the United Kingdom, Switzerland and India. Flash TTS is rolling out through the Gemini API, Google AI Studio and Gemini Notebook, while Flash-Lite is rolling out through the Gemini API and Google AI Studio, with enterprise API access listed as coming soon and no pricing published. Google launches Gemini 3.8 voice models with Hume founder Alan Cowen credited The former Hume AI CEO is credited on Google's release of models for custom voice design and line-by-line performance control, with 30-second voice replication built in. By RuntimeWire Staff https://runtimewire.com/author/runtimewire-staff ยท Published Primary source: Google https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-text-to-speech/ Why it matters Google is pairing custom voice creation with existing developer and workplace distribution. Cowen's move from Hume AI puts a founder of a specialist voice startup inside that effort. Pricing, latency and the practical enforcement of voice consent remain key tests for production use. Alan Cowen, the former Hume AI founder and a Google DeepMind Director of Research Science, is credited alongside Leland Rechis on Google's https://google.com/?ref=runtimewire launch of Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-text-to-speech/?ref=runtimewire on September 23rd. The models add custom voice design and line-by-line direction to Google's audio products. Cowen's career gives the announcement a clear throughline. His research biography https://www.alancowen.com/bio?ref=runtimewire describes work on how emotional behavior can be measured, predicted and modeled. Google credits him on a release whose central pitch is that developers can design a voice's character and direct its performance, rather than simply select a preset. From a voice preset to a performance Google says Flash TTS is aimed at creating original character voices for games, audiobooks, podcasts and interactive media. A user can describe a voice in natural language, then direct delivery line by line with cues for pacing, acting, dialect and conversational sounds such as laughter or a short acknowledgment. The model can also stage a two-person exchange from one script and maintain character voices through long-form audio, according to Google's announcement https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-text-to-speech/?ref=runtimewire . Flash-Lite is positioned for high-volume work such as dubbing, audio production and voice agents. Google calls it cost-efficient, but did not publish a price or a cost comparison in the announcement. That leaves developers unable to judge the economic difference between the two models from the launch materials alone. Google says the models support more than 100 languages and dialects, offer a library of more than 2,000 production-ready voices, and can replicate a voice from a 30-second audio sample. Those are Google's product claims, detailed in its launch announcement https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-text-to-speech/?ref=runtimewire . Google says voice replication includes built-in consent verification. It also says generated audio carries SynthID, its audio watermark, and C2PA credentials. Voice replication is unavailable in Illinois, Texas, the European Economic Area, the United Kingdom, Switzerland and India. The safeguards answer part of the problem; they do not establish how Google handles disputed consent or enforces rights claims after a voice has been created. The announcement does not detail those procedures. Google also says its safeguards apply to replication, while watermarking applies across generated Gemini Audio clips. Distribution is part of the product Google is distributing the models through existing products rather than launching a separate voice service. Flash TTS is rolling out to developers through the Gemini API and Google AI Studio, and to consumers through Gemini Notebook https://notebook.google.com/?ref=runtimewire . Enterprise access through API is listed as coming soon. Flash-Lite began rolling out through the Gemini API and Google AI Studio and is available in Google Vids http://vids.new/?ref=runtimewire ; access through the Gemini Enterprise API is listed as coming soon, according to Google's announcement https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-text-to-speech/?ref=runtimewire . That placement gives Google a route into developer, enterprise and creator workflows already using its tools. It also puts the models beside existing Gemini audio capabilities: Google introduced Gemini 3.8 Live and Gemini 3.5 Transcribe on September 15th, describing those releases as tools for real-time voice applications. The new TTS models address a different task: producing directed speech for a script, a character or an agent's response. Google's earlier Gemini 3.1 Flash TTS release https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-1-flash-tts/?ref=runtimewire emphasized expressive speech and audio tags for controlling style and pacing. The 3.8 announcement pushes further into voice identity and performance direction, while keeping a lower-cost, higher-volume option in the lineup. That is a product strategy as much as a model update: offer creative control for authors and media makers, then offer a scale-oriented model for companies that need to generate more speech. Cowen's old thesis meets Google's distribution Cowen's move from Hume to Google DeepMind was reported in January as part of a licensing deal that also brought Hume engineers to Google. WIRED reported https://www.wired.com/story/google-hires-hume-ai-ceo-licensing-deal-gemini/?ref=runtimewire that Hume would continue supplying its technology to other AI labs. His credit on this launch places a founder of a specialist voice company inside Google's effort to build speech products across its model and software portfolio. The overlap is unusually direct: Google's own announcement evaluates Flash TTS using Hume AI's Voice Design Benchmark https://www.hume.ai/rw-voice-eq?ref=runtimewire and Overall Quality Index, and reports top placements for the new models. Google also cites blind preference evaluations on Voice Arena. Those results are reported by Google; they do not, by themselves, settle how the products compare for a particular use case or what they cost to run. The commercial incentive is visible in the market's recent numbers. ElevenLabs reported it passed $500 million in annual recurring revenue in the first four months of 2026, after ending 2025 at $350 million https://elevenlabs.io/blog/500m-arr-and-new-investors?ref=runtimewire . Those figures are company-reported, but they show why expressive speech and voice agents are drawing investment and product attention. Google's bet is that it can put comparable voice creation tools inside places developers and businesses already build and work. The announcement supplies no pricing, latency targets or product-specific usage figures. Google's label of these as its most expressive audio models is also a company claim, not a result independently established by the announcement. The immediate test for customers will be whether the finer directing controls and existing distribution make the models useful in production, especially where voice rights, consistency and per-minute cost matter as much as how a sample sounds.