cd /news/large-language-models/google-tests-new-gemini-4-pro-checkp… · home topics large-language-models article
[ARTICLE · art-137948] src=testingcatalog.com ↗ pub= topic=large-language-models verified=true sentiment=↑ positive

Google tests new Gemini 4 Pro checkpoints, early outputs

Google has begun internal testing of early Gemini 4 Pro checkpoints under the code name "argon," according to a leak from Lentils, with developer Bee posting a demo on X of a monographic website the model generated in fourteen minutes. Bee said the output was "so good" and that Google had done a great job with the design, addressing Gemini's long-standing weakness in frontend design taste. Google is also testing Projects for the Gemini desktop app, which will work like the workspaces in ChatGPT or Claude with dedicated hubs for reference files and custom instructions.

by read3 min views2 publishedSep 23, 2026
Google tests new Gemini 4 Pro checkpoints, early outputs
Image: Testingcatalog (auto-discovered)

Google has been steadily dropping newer Flash models over the past few months, but hardcore users have only been waiting for one thing: a new frontier-class model. The recent Gemini 3.8 Flash release brought some interesting coding improvements, but it would be a disservice to call it a breakthrough model.

Gemini models have long had a reputation for producing awkward, unoriginal web designs when asked to create frontend code. That might finally be on its way out, if early test results of what appears to be Gemini 4 Pro are anything to go by.

The leak from Lentils was among the first to mention it a few days ago, saying Google had begun testing early checkpoints of Gemini 4 Pro within the company under the code name "argon."

Since then, others have also shared their early Gemini 4 Pro test results on X. A demonstration posted by developer Bee on X featuring a fully monographic website that looked extremely clean was pretty impressive.

Although the model took fourteen minutes to produce the page, the final result had nothing whatever in common with the robotic templates which we typically receive from AI tools. Bee said that he was surprised at how clean and well-polished the monographic website had become "it looks so good", and mentioned that Google had done a great job with the design, the earlier problems regarding design taste now appearing to have been resolved.

The audio effects used to create the sound of sketching on canvas paper are also a neat touch that adds to the interface's tactile feel.

Another AI benchmarking account also shared Gemini 4 Pro's output in an attempt to create an Xbox controller SVG. We've seen similar tests from Gemini models in the past, but none have had the same level of polish as this one. So it's almost certain this is Gemini 4 Pro's work of art.

Way back in July, Logan Kilpatrick hinted that the team is cooking with "Gemini 3.5 Pro," but it looks like Google may not have been too excited about the results at the time. With all the progress made since then, it only makes sense that the team would skip the 3.5 moniker and go straight to 4 Pro. Plus, with these early results, it might be worth the big jump instead of a point release.

That said, while Gemini 4.0 Pro is still in the oven, Google is also working on enhancements for the Gemini desktop app, which we spotted in testing. In short, Projects will work a lot like the workspaces in ChatGPT or Claude. You’ll have dedicated hubs for reference files and custom instructions, so each new conversation uses that shared context instead of making you start over every time.

These tests also come at an interesting time, as we're seeing some of the first signs of what could be Anthropic's upcoming Fable 5.2 model. So it'll be interesting to see what Google has been working on for the past few months and how it compares to Anthropic's next frontier model.

The results from these current tests suggest that Gemini 4 Pro could possibly have a real chance of becoming a worthy frontier-class model from Google. But we'll have to wait and see how it fares in head-to-head comparisons with Astra, Grok 4.7, and, of course, Fable 5.2 to see how it stacks up against the competition.

── more in #large-language-models 4 stories · sorted by recency
── more on @google 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/google-tests-new-gem…] indexed:0 read:3min 2026-09-23 ·