cd /news/large-language-models/gemma-4-gets-a-stealth-update-that-f… · home topics large-language-models article
[ARTICLE · art-61777] src=the-decoder.com ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Gemma 4 gets a stealth update that fixes tool calling bugs and truncated responses under the same name

Google shipped a stealth update to its open AI model Gemma 4 that fixes tool calling bugs and truncated responses, while speeding up performance on Nvidia Hopper GPUs by 25 to 70 percent with Flash Attention 4 enabled. The update also reduces time to first token by up to 31 percent and allows users to raise the max_soft_tokens parameter from 280 to 1,120 for sharper OCR results. The community criticized Google for releasing the update under the same name instead of a separate version like Gemma 4.1.

read1 min views54 publishedJul 16, 2026
Gemma 4 gets a stealth update that fixes tool calling bugs and truncated responses under the same name
Image: The Decoder

**Google shipped an update to its open AI model Gemma 4 that speeds up performance on Nvidia Hopper GPUs, fixes tool calling bugs, and addresses problems with truncated responses. **Turning on Flash Attention 4 boosts the speed at which the model processes incoming prompts by 25 to 70 percent, according to Google. Time to first token drops by up to 31 percent. Google also fixed bugs in tool calling, the feature that lets the model trigger external tools on its own.

Google says it also cut down on cases where the model would cut answers short or return incomplete responses. For image processing, users can manually raise the "max_soft_tokens" parameter from 280 to 1,120 to get sharper OCR results and support resolutions up to 2.51 megapixels. Google put up an interactive configurator on Hugging Face for that. The published benchmarks only compare the 31B and E4B variants against their predecessors, but the Hugging Face repository shows that all parameter sizes in this model generation got updated, including the newest 12B release. The community has pushed back on Google shipping the update under the same "Gemma 4" name instead of tagging it as a separate version like "Gemma 4.1."

AI News Without the Hype – Curated by Humans

					Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.				

					Subscribe now

X/Google

── more in #large-language-models 4 stories · sorted by recency
── more on @google 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/gemma-4-gets-a-steal…] indexed:0 read:1min 2026-07-16 ·