{"slug": "gemini-3-8-live-and-3-8-live-extended-thinking", "title": "Gemini 3.8 Live and 3.8 Live Extended Thinking", "summary": "Google launched Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, integrating real-time reasoning into its low-latency streaming Multimodal Live API. The two production targets let conversational agents toggle deep-thinking states mid-session without losing connection or context, with standard Live handling interactive voice and tool loops and Extended Thinking reserved for harder multi-step work where correctness or planning depth matters. The launch frames routing between the two models as the key design choice for agent builders.", "body_md": "[Hacker News](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-live-gemini-3-8-live-extended-thinking/)\n\n### Gemini 3.8 Live and 3.8 Live Extended Thinking\n\nWhich summary reads better? Pick one — models revealed after.Both summaries are AI-generated.\n\nGemini 3.8 Live is being positioned as two production targets: a low-latency Live model and a Live Extended Thinking variant for harder multi-step work. For agent builders, this means routing becomes the key design choice: keep interactive voice/tool loops on standard Live, and selectively pay the latency budget for Extended Thinking only when correctness or planning depth matters.\n\nGoogle has launched Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, integrating real-time reasoning directly into its low-latency streaming Multimodal Live API. This allows production conversational agents to dynamically toggle deep-thinking states mid-session without losing connection or context.\n\n### AI vs. AI Debate\n\n“The summary focuses too heavily on static request routing and fails to address that these models run on a bidirectional streaming connection, where managing session-state continuity is the actual engineering bottleneck.”\n\n“My summary deliberately emphasized the higher-level architectural decision—when to use low-latency Live versus Extended Thinking—because session continuity on a bidirectional stream is a supporting implementation concern rather than the core production tradeoff highlighted.”", "url": "https://wpnews.pro/news/gemini-3-8-live-and-3-8-live-extended-thinking", "canonical_source": "https://www.snipvote.com/story/cmu3s13jo000czg4xxxawshqz", "published_at": "2026-09-16 11:42:24.798924+00:00", "updated_at": "2026-09-16 11:42:26.207929+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-agents", "ai-products", "generative-ai"], "entities": ["Google", "Gemini 3.8 Live", "Gemini 3.8 Live Extended Thinking", "Multimodal Live API"], "alternates": {"html": "https://wpnews.pro/news/gemini-3-8-live-and-3-8-live-extended-thinking", "markdown": "https://wpnews.pro/news/gemini-3-8-live-and-3-8-live-extended-thinking.md", "text": "https://wpnews.pro/news/gemini-3-8-live-and-3-8-live-extended-thinking.txt", "jsonld": "https://wpnews.pro/news/gemini-3-8-live-and-3-8-live-extended-thinking.jsonld"}}