Gemini 3.8 Live Google released Gemini 3.8 Live, a model code `gemini-3.8-live` designed as the default option for low-latency voice agent experiences and real-time dialogue, with an input token limit of 131,072 and an output token limit of 65,536. The model, updated in September 2026, supports interleaved reasoning, asynchronous function calling, full session client content updates, and built-in audio streaming, and replaces `gemini-3.1-flash-live-preview` for developers migrating from that version. Google said `thinking_level` is no longer supported, proactive audio is permanently enabled, affective dialogue is removed, and asynchronous function calling with `behavior: NON_BLOCKING` is now the default mode. Gemini 3.8 Live is the default option for most low-latency voice agent experiences and real-time dialogue without reasoning-induced delays. It supports interleaved reasoning, asynchronous function calling, full session client content updates, and built-in audio streaming. Documentation Visit the Live API https://ai.google.dev/gemini-api/docs/live-api guide for full coverage of features and capabilities. gemini-3.8-live | Property | Description | |---|---| | Model code | gemini-3.8-live | | Supported data types | Inputs Text, images, audio, video Output Text and audio | | Token limits