Gemini 3.8 Live Extended Thinking audio-to-audio Google released Gemini 3.8 Live Extended Thinking, an audio-to-audio model that performs background reasoning and asynchronous tool calls while streaming continuous audio, with a 131,072-token input limit and 65,536-token output limit. The model, updated September 2026, requires clients to keep listening after turnComplete: true and to track the interaction_status field (IN_PROGRESS or IDLE), and it supports only asynchronous non-blocking function calling. Background reasoning is configured via thinking_config with thinking_level set to low, medium, or high; MINIMAL is not supported. Gemini 3.8 Live Extended Thinking is our high-reasoning audio-to-audio model recommended when higher background reasoning is required for complex, multi-step problem solving during real-time voice interactions. It processes background reasoning and asynchronous tool calls while streaming continuous audio responses. Documentation Visit the Live API https://ai.google.dev/gemini-api/docs/live-api guide for full coverage of features and capabilities. gemini-3.8-live-extended-thinking | Property | Description | |---|---| | Model code | gemini-3.8-live-extended-thinking | | Supported data types | Inputs Text, images, audio, video Output Text and audio | | Token limits