Google DeepMind Unveils Gemini 4 Argon with 1M Output Tokens for Coding, Knowledge Work and Cyber Defense Google DeepMind announced Gemini 4 Argon, its first Gemini 4-generation frontier model, which generates up to 1M output tokens in a single response versus 64K on earlier Gemini models and launches at introductory pricing of $2 per 1M input tokens and $10 per 1M output tokens. Google said Argon leads outright on 12 of 18 benchmarks and ties for first on 1, including a state-of-the-art 77.9% on DeepSWE v1.1 and 91.7% on LVBench, while trailing GPT-6 Astra on FrontierSWE v2 (55.0% vs 65.5%) and Claude Opus 5.5 on Terminal-Bench 4.0 (57.4% vs 66.4%). Wiz is already using Argon through its Scan for Good initiative, where the model found a critical vulnerability in healthcare software used by hospitals worldwide that Google says previous frontier models missed. Google DeepMind has just announced Gemini 4 Argon https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/ , its new frontier model and the first model of the Gemini 4 generation. It targets long-horizon software engineering, enterprise knowledge work in legal and finance, and cybersecurity defense. The biggest technical change is output length. Argon can generate up to 1M tokens in a single response, up from 64K on earlier Gemini models. What Google Announced Google DeepMind https://x.com/GoogleDeepMind/status/2105388084154056939 described Argon as built for complex workflows across coding, enterprise knowledge work and cybersecurity defense. Google is taking a phased approach. It is participating in the U.S. government’s voluntary process for pre-release model access. It will gather feedback from early testers and iterate on guardrails before a wider release. Pricing is already public. Argon launches at an introductory $2 per 1M input tokens and $10 per 1M output tokens. Cached input tokens get a 95% discount, which works out to $0.10 per 1M. After the introductory period, pricing moves to $4 input and $20 output. Logan Kilpatrick https://x.com/OfficialLoganK/status/2105388054274080946 confirmed the introductory $2 in and $10 out pricing. Why the 1M Output Limit Matters Current frontier APIs cap a single response far lower. Claude Opus 5.5 https://venturebeat.com/technology/anthropic-releases-claude-opus-5-5-beating-fable-5-1-on-key-agentic-benchmarks-at-60-cheaper-api-price , Claude Fable 5.1 https://platform.claude.com/docs/en/models/fable-5-1/overview and GPT-6 Astra https://openrouter.ai/openai/gpt-6-astra each allow 128K output tokens. Google team states that Argon can think deeply and generate hundreds of thousands of tokens in one trajectory. For developers, that means large refactors or long reports without splitting work across turns. The cost is real, though. A full 1M output tokens costs $10 at introductory pricing and $20 after. Google has not disclosed Argon’s input context window. Benchmarks: Where Argon Leads and Where It Trails Google compared Argon against GPT-6 Astra, Claude Opus 5.5 and Claude Fable 5.1. Argon leads outright on 12 of 18 benchmarks and ties for first on 1. Where it leads: - DeepSWE v1.1 long-horizon software engineering : 77.9%, a new state of the art. Opus 5.5 scores 74.2% and GPT-6 Astra 74.1%. - Vals Index https://www.vals.ai/benchmarks/vals index economic impact across finance, coding, legal and tax : 68.9%, ranked first. - AutomationBench Zapier, end-to-end business execution : 51.3%, ranked first. Opus 5.5 scores 42.5%. - Harvey Legal Agent Benchmark: 19.6%, against 5.4% for GPT-6 Astra. - LVBench long video understanding : 91.7%, a new state of the art. Where it trails: - FrontierSWE v2: 55.0%, behind GPT-6 Astra at 65.5%. - Terminal-Bench 4.0: 57.4%, behind Claude Opus 5.5 at 66.4%. - OSWorld-2.0 computer use : 69.2%, behind GPT-6 Astra at 72.6%. Artificial Analysis https://x.com/ArtificialAnlys/status/2105392625788637299 reported that Argon equals GPT-6 Astra on its Intelligence Index at 60% of the cost per task, using discounted prices. Cyber Defense: Find, Validate, Patch Google trained Argon to autonomously find, validate and patch critical software vulnerabilities. Trusted defenders and internal Google teams receive it without cyber guardrails. On CWE-bench v1 https://cwe-bench.com/ , which tests vulnerability remediation, Argon ties for first at 68%. The rival models on that leaderboard run inside their own agent harnesses. Wiz https://www.wiz.io/ is already using Argon through its Scan for Good https://www.wiz.io/scan-for-good initiative. The model found a critical vulnerability in healthcare software used by hospitals worldwide. Google says previous frontier models had missed it. Before broad release, Google is strengthening safeguards in 4 areas: - Misuse defenses for cyber and CBRN risks, including activation monitoring, under its Frontier Safety Framework https://deepmind.google/blog/strengthening-our-frontier-safety-framework/ . - Indirect prompt injection resistance, where Argon leads Gray Swan’s IPI benchmark. - Misalignment monitoring of chain-of-thought and actions, with the ability to stop execution. - Sealed, isolated sandboxes for high-risk training and evaluations. Argon Inside Google Thousands of Googlers already use Argon. Google shared 4 internal results: - Argon agents applied memory optimizations across data centers, freeing over 300 TiB, with 500 TiB to 1 PiB projected. - Agents replaced 32K lines of SIMD code in the libgav1 Rust port. The decoder runs 2.7x faster with identical output. - Agents are migrating C/C++ codebases to Rust, up to 800K+ lines in the Fuchsia Zircon kernel. - Argon beat a published quantum algorithm baseline by 40% in minutes. Comparison: Gemini 4 Argon vs Closest Competitors | Feature | Gemini 4 Argon | Claude Opus 5.5 | Claude Fable 5.1 | GPT-6 Astra | |---|---|---|---|---| | Developer | Google DeepMind | Anthropic | Anthropic | OpenAI | | Availability | Fairwind Program only | Claude API and clouds | Claude API and clouds | OpenAI API | | Max output per response | 1M tokens | 128K | 128K | 128K | | Context window | Not disclosed | 1M | 1M | 1.05M | | Input / output price per 1M | $2 / $10 intro, then $4 / $20 | $4 / $20 | $10 / $50 | $10 / $50 | | Cached input per 1M | $0.10 intro | $0.20 | $0.25 | $1.00 | | Open weights | No | No | No | No | | DeepSWE v1.1 | 77.9% | 74.2% | 67.4% | 74.1% | | Vals Index | 68.9% | 67.0% | 65.8% | 63.1% | | FrontierSWE v2 | 55.0% | 62.3% | 56.3% | 65.5% | | Terminal-Bench 4.0 | 57.4% | 66.4% | 57.9% | 58.2% | | CWE-bench v1 | 68% tie | 67% | 58% | 68% tie | Sources: Google https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/ , Anthropic Opus pricing https://www.anthropic.com/claude/opus , Anthropic Fable 5.1 docs https://platform.claude.com/docs/en/models/fable-5-1/overview , OpenAI GPT-6 Astra docs https://developers.openai.com/api/docs/models/gpt-6-astra , OpenRouter https://openrouter.ai/openai/gpt-6-astra . Benchmark scores are from Google’s published comparison. GPT-6 Astra prices are its short-context tier. Key Takeaways - Gemini 4 Argon raises the output limit from 64K to 1M tokens. - It leads DeepSWE v1.1 77.9% and the Vals Index 68.9% . - It trails on FrontierSWE v2, Terminal-Bench 4.0 and OSWorld-2.0. - Introductory pricing of $2 / $10 is half of Claude Opus 5.5. - Access is limited to Fairwind cyber defenders; no public release date yet. Check out the technical details https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/ . All credit goes to the researcher of this project. Also, feel free to follow us on Twitter https://x.com/intent/follow?screen name=marktechpost and don’t forget to join our 150k+ML SubReddit https://www.reddit.com/r/machinelearningnews/ and Subscribe to our Newsletter https://magic.beehiiv.com/v1/f5e63dd4-5653-4f09-83e2-321a8b1ba526?email={{email}} . Wait are you on telegram? now you can join us on telegram as well. https://t.me/machinelearningresearchnews Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us https://forms.gle/MJjjVDPS7whH8Ngs6 Asif Razzaq is the CEO of Marktechpost AI Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.