cd /news/artificial-intelligence/gemini-3-8-live-tops-voice-benchmark… · home topics artificial-intelligence article
[ARTICLE · art-131378] src=vibeleaderboard.ai ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

Gemini 3.8 Live tops voice benchmarks while calling tools mid-chat

Google shipped Gemini 3.8 Live and an Extended Thinking variant that rank first on speech-to-speech and voice-agent benchmarks, and the models can now call tools mid-conversation without breaking the dialogue. The release also coincides with GPT-5.5 leaving ChatGPT and default Codex on October 14, though it remains available through the API Platform and in Codex sessions authenticated with an API key. In a separate finding, EvolveScaler reports that frontier models' accuracy collapses to 11.3 percent on its hardest tier when tracking event logs with retracted and corrected records.

read1 min views1 publishedSep 16, 2026

Google shipped Gemini 3.8 Live and an Extended Thinking variant that ranks first on speech-to-speech and voice-agent benchmarks, and it can now call tools without breaking the conversation. Read: Google shipped Gemini 3.8 Live and an Extended Thinking variant that ranks first on speech-to-speech and voice-agent benchmarks, and it can now call tools without breaking the conversation. Read: GPT-5.5 leaves ChatGPT and default Codex on October 14, but stays available through the API Platform and in Codex sessions authenticated with an API key. Read: Strix's autonomous pentesting agent chained a public container registry to a live GitHub admin token on Baseten's production repos, unattended, in under half an hour. Read: A new "Disallow AI Training" toggle lets sites opt out of training use without losing search visibility, alongside an "Accountable" crawler standard shared with Apple, Google, and Microsoft. Read: SemiAnalysis argues local moratoriums are too small relative to planned and under-construction capacity to explain any real slowdown in the US AI datacenter buildout. Read: EvolveScaler tests whether models can track event logs where records get retracted and corrected, and frontier models' accuracy collapses to 11.3 percent on its hardest tier. Read: Anthropic's new beta integration pulls Salesforce accounts, opportunities, and pipeline data into Claude for call prep, deal review, and forecasting.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @google 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/gemini-3-8-live-tops…] indexed:0 read:1min 2026-09-16 ·