{"slug": "gemini-3-8-live-tops-voice-benchmarks-while-calling-tools-mid-chat", "title": "Gemini 3.8 Live tops voice benchmarks while calling tools mid-chat", "summary": "Google shipped Gemini 3.8 Live and an Extended Thinking variant that rank first on speech-to-speech and voice-agent benchmarks, and the models can now call tools mid-conversation without breaking the dialogue. The release also coincides with GPT-5.5 leaving ChatGPT and default Codex on October 14, though it remains available through the API Platform and in Codex sessions authenticated with an API key. In a separate finding, EvolveScaler reports that frontier models' accuracy collapses to 11.3 percent on its hardest tier when tracking event logs with retracted and corrected records.", "body_md": "Google shipped Gemini 3.8 Live and an Extended Thinking variant that ranks first on speech-to-speech and voice-agent benchmarks, and it can now call tools without breaking the conversation.\nRead: Google shipped Gemini 3.8 Live and an Extended Thinking variant that ranks first on speech-to-speech and voice-agent benchmarks, and it can now call tools without breaking the conversation.\nRead: GPT-5.5 leaves ChatGPT and default Codex on October 14, but stays available through the API Platform and in Codex sessions authenticated with an API key.\nRead: Strix's autonomous pentesting agent chained a public container registry to a live GitHub admin token on Baseten's production repos, unattended, in under half an hour.\nRead: A new \"Disallow AI Training\" toggle lets sites opt out of training use without losing search visibility, alongside an \"Accountable\" crawler standard shared with Apple, Google, and Microsoft.\nRead: SemiAnalysis argues local moratoriums are too small relative to planned and under-construction capacity to explain any real slowdown in the US AI datacenter buildout.\nRead: EvolveScaler tests whether models can track event logs where records get retracted and corrected, and frontier models' accuracy collapses to 11.3 percent on its hardest tier.\nRead: Anthropic's new beta integration pulls Salesforce accounts, opportunities, and pipeline data into Claude for call prep, deal review, and forecasting.", "url": "https://wpnews.pro/news/gemini-3-8-live-tops-voice-benchmarks-while-calling-tools-mid-chat", "canonical_source": "https://www.vibeleaderboard.ai/intel/brief/2026-09-16", "published_at": "2026-09-16 11:15:46+00:00", "updated_at": "2026-09-16 12:13:32.729400+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-products", "ai-agents", "ai-research"], "entities": ["Google", "Gemini 3.8 Live", "Extended Thinking", "GPT-5.5", "ChatGPT", "Codex", "EvolveScaler", "Anthropic"], "alternates": {"html": "https://wpnews.pro/news/gemini-3-8-live-tops-voice-benchmarks-while-calling-tools-mid-chat", "markdown": "https://wpnews.pro/news/gemini-3-8-live-tops-voice-benchmarks-while-calling-tools-mid-chat.md", "text": "https://wpnews.pro/news/gemini-3-8-live-tops-voice-benchmarks-while-calling-tools-mid-chat.txt", "jsonld": "https://wpnews.pro/news/gemini-3-8-live-tops-voice-benchmarks-while-calling-tools-mid-chat.jsonld"}}