# Google unveils Gemini 4 Argon with SOTA score on DeepSWE

> Source: <https://www.testingcatalog.com/google-unveils-gemini-4-argon-with-sota-score-on-deepswe/>
> Published: 2026-09-30 20:33:39+00:00

Google has unveiled [Gemini 4 Argon](https://www.testingcatalog.com/gemini-4-pro-frontend-ui-taste-leak/), a frontier model built to sustain deep reasoning across long, complex workflows in software engineering, enterprise knowledge work and cybersecurity defense. Access is initially limited to trusted cyber defenders through Google’s Fairwind Program while the company gathers feedback and develops its guardrails. Google is also participating in the U.S. government’s voluntary process for pre-release model access. Wider access is planned as soon as possible for developers, enterprises and consumers, starting with paid API customers and Google AI Ultra subscribers.

Argon will launch at an introductory price of $2 per million input tokens and $10 per million output tokens, with cached input tokens priced at a 95% discount to the input rate. After the introductory period, pricing will rise to $4 per million input tokens and $20 per million output tokens. Google is also lifting the model’s output limit from 64K to 1 million tokens, giving it room to reason and generate hundreds of thousands of tokens within one trajectory.

Thousands of Google employees are already using Argon for coding, research and writing. Google says the model beat a published quantum algorithm optimization baseline by 40% within minutes, while a group of Argon agents found data-center memory optimizations that freed more than 300 TiB after rollout, with total projected savings of 500 TiB to 1 PiB. Agents are also supporting C and C++ migrations to Rust, from core libraries with tens of thousands of lines to the Fuchsia Zircon kernel at more than 800K lines. Because these systems are critical, the migrations undergo automated and manual auditing, emulation testing and review before production. In another project, Argon replaced 32K lines of SIMD code in a Rust port of the libgav1 video decoder, producing safe Rust that ran 2.7 times faster with identical output.

The model scored 77.9% on DeepSWE v1.1 for long-horizon software engineering, ranked first on the Vals Index and scored 51.3% on Zapier’s AutomationBench. It also reached 91.7% on the long-video LVBench test. In cybersecurity, Argon tied for first on CWE-bench v1 with 68%. Google says Wiz used Argon through its Scan for Good program to uncover a critical vulnerability that exposed sensitive personal information in healthcare software used by hospitals worldwide, a risk missed by earlier frontier models.

Trusted defenders and [Google’s](https://www.testingcatalog.com/tag/google/) internal teams will receive Argon without cyber guardrails for defensive work. Broader deployment remains phased while Google strengthens protections against cyber and CBRN misuse, indirect prompt injection and misaligned agent behavior, alongside hardened sandbox environments for high-risk training and evaluations.
