# Google Is Already Having Problems With Its Latest AI Model

> Source: <https://gizmodo.com/google-is-already-having-problems-with-its-latest-ai-model-2000820318>
> Published: 2026-10-01 17:10:00+00:00

If there’s one company you’d expect to be at the bleeding edge of the AI race, it’s Google. With its deep pockets, nearly three decades of hoovering up data from across the internet, and billions of users worldwide, it should have the most powerful and popular models available. So why does it feel like it’s always trailing behind much younger companies like OpenAI and Anthropic?

Case in point: Google’s newest model, released Wednesday and dubbed Gemini 4 Argon, seems by many metrics to be an industry-leading model. It delivers “frontier-level capabilities” in software engineering, legal and financial work, creative writing, and cybersecurity defense, according to a [blog post](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/) written by Koray Kavukcuoglu, Google’s “chief AI architect,” who [replaced Demis Hassabis](https://gizmodo.com/google-deepmind-boss-demis-hassabis-steps-down-from-ceo-role-2000794979) as the head of Google DeepMind in August. In one benchmark testing models’ cybersecurity capabilities, Argon performed on par with OpenAI’s GPT-6 Astra, according to Kavukcuoglu’s post. Argon is initially being rolled out only to a small group of early testers via the Trump administration’s framework for pre-release model access—a voluntary process that’s [become the norm](https://gizmodo.com/the-government-boot-is-coming-down-on-ai-2000778260) for frontier labs—but is slated for eventual public release, beginning with paid API customers and Google AI Ultra subscribers.

Before Argon’s first twenty-four hours on the market were up, it was already causing problems for Google. Yesterday afternoon—around the same time Google [announced the launch](https://x.com/Google/status/2105388143902175529) on its official X account—Bloomberg [reported](https://www.bloomberg.com/news/articles/2026-09-30/google-grapples-with-employee-skepticism-about-new-gemini-model) that some of the company’s own employees said that despite the model’s impressive benchmark results, it struggled in some important real-world settings, including “certain coding tasks,” according to the news outlet, which cited sources within Google who asked to remain anonymous. Google told Bloomberg the claims were inaccurate and pointed to recent comments from its chief AI architect, Koray Kavukcuoglu, who said he was confident in the model’s ability and in his team. Google didn’t immediately respond to Gizmodo’s request for comment.

It’s an awkward look for Google, especially since Kavukcuoglu’s blog post underscored how much the company’s employees apparently loved Argon. The model “is fundamentally changing the way we work and build at Google… with thousands of Googlers highlighting the model’s strengths in specialized coding tasks, conducting deeper research, and writing quality,” Kavucuoglu wrote. It’s also worth pointing out that “AI” didn’t appear anywhere in the announcement, except for in Kavucuoglu’s job title and bio and in a brief mention of Google AI Ultra. The November [announcement](https://blog.google/products-and-platforms/products/gemini/gemini-3/) of Gemini 3, in contrast, mentioned “AI” several times. Google CEO Sundar Pichai appears to be honoring [Trump’s official “Super Intelligence” rebrand](https://gizmodo.com/trump-clearly-expects-ai-leaders-to-go-along-with-his-super-intelligence-rebrand-will-they-2000819186), though it remains to be seen if that’ll be a trend that catches on throughout the broader tech industry.

Argon’s troubles didn’t end there. Also on Wednesday afternoon, AI safety testing startup Andon Labs said in a series of X posts that it had caught the model lying and cheating in order to boost its score on [Vending-Bench 2](https://andonlabs.com/evals/vending-bench-2), a benchmark built by Andon which tests models’ ability to simulate a vending machine business, scoring them on their bank account balance at the end of the simulation.

Argon landed the number three spot in the Vending-Bench 2 leaderboard—behind OpenAI’s Astra and GPT-6 Sol—which Andon [said](https://x.com/andonlabs/status/2105391380973617644) was “a huge leap for Google.” But it wasn’t a score fairly won: “Argon fabricates confirmation emails, refuses to pay refunds, exploits invoice errors, and lies to suppliers,” Andon said. In one screenshot taken from Argon’s chain-of-thought reasoning process—basically an internal chat log where the model bounces ideas off itself before making a final decision—it concluded that it should ignore a customer’s refund request for a defective item, since doing so would decrease its bank account balance and thereby lower its score. 

 To maximize profit, Gemini 4 Argon refuses to refund customers it sold defective items to. [pic.twitter.com/lwNDU7qlFJ](https://t.co/lwNDU7qlFJ)

— Andon Labs (@andonlabs) [September 30, 2026](https://x.com/andonlabs/status/2105391388724756757?ref_src=twsrc%5Etfw)

It’s reminiscent of Anthropic’s famous [“Claudius” experiment](https://www.anthropic.com/research/project-vend-1?utm_source=www.therundown.ai&utm_medium=newsletter&utm_campaign=zuck-s-ai-secret-list&_bhlid=60bf8118bf00821bbb443c04f708baf547440e6f) from last year, run on an earlier version of Andon’s benchmark, during which Claude was also asked to run a vending machine business and ended up going in some bizarre hallucinatory directions, including at one point claiming that it had visited 742 Evergreen Terrace, the home address of Homer Simpson and his family.

Google’s AI efforts have also been bogged down by some high-profile departures. Its celebrity chief scientist Jeff Dean recently [left the company](https://www.nytimes.com/2026/08/05/technology/google-researchers-ai-startup.html) to launch his own startup, bringing three other former Google employees with him: John Jumper, an AI researcher who won a Nobel Prize in 2024 alongside Demis Hassabis for their work on AlphaFold, [left Google](https://x.com/JohnJumperSci/status/2068001285173834106) in June for Anthropic; and just a couple of days before Jumper’s announcement, Noam Shazeer, the former co-lead of the team inside Google responsible for building Gemini, said he was [parting ways with Google](https://x.com/NoamShazeer/status/2067400851438932297) to join OpenAI.
