[AI in Practice] Gemini 3.8 Flash TTS Launch: I built a "Learn Japanese with MVs" Web App and burned through my daily quota. A developer built a Japanese-language learning web app that uses Google's newly generally available Gemini 3.8 Flash TTS and Flash-Lite TTS models to generate sentence-by-sentence pronunciation audio, hitting the Tier 1 quota of 100 requests per day. The writeup details the new interactions and voices API structure, custom voice limits of 200 per project, WAV 16-bit PCM mono 24 kHz output, and regional restrictions on voice replication, while noting the app's lyrics-sourcing approach raises copyright considerations. Every time I see a new Gemini feature, my first thought is "Can I connect it to my LINE Bot?" The Gemini API release notes https://ai.google.dev/gemini-api/docs/changelog from 9/22 stated that Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS are officially launched GA , and the official blog simultaneously posted Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-text-to-speech/ . As usual, I listed a bunch of LINE Bot ideas: bedtime stories in parents' voices, turning group chats into radio dramas, morning dual-host podcasts... Halfway through the list, I realized that what actually struck me most wasn't a bot, but the fact that " directing tone sentence-by-sentence " is particularly suitable for teaching pronunciation. So the topic changed to a Web App: As it turned out, there was nothing to complain about regarding the quality of the TTS itself; what really took time were the things surrounding it that weren't in the documentation. Two models were released in this GA: | Model | Positioning | |---|---| | gemini-3.8-flash-tts | Flagship model, emphasizes vocal performance and character creation, allows sentence-by-sentence performance control | | gemini-3.8-flash-lite-tts | Cheap, fast, suitable for high-volume generation | Compared to the previous generation, there are three things I find truly useful: voice id for repeated use. style e.g., "speak slowly, pronounce every syllable clearly" , and it supports tags like