Google announced Gemini 4 Argon on September 30, roughly ten months after Gemini 3. On Google’s evaluations, Argon leads or ties GPT-6 Astra and Claude Opus 5.5 on 14 of 19 benchmarks. Artificial Analysis scores it at 53, level with Astra and below Opus 5.5 at 58 and Sonnet 5.5 at 56. Google is back in frontier competition, though it still trails OpenAI and Anthropic overall.
Introductory pricing is $2 per million input tokens and $10 per million output tokens, rising to $4/$20 at the standard rate. We use this input/output convention throughout. Access is initially limited to trusted testers, with paid API customers and Google AI Ultra subscribers next in line.
Two months ago, many investors thought Google was stepping back from frontier models. Gemini 3.5 Pro had missed its launch window and been shelved, leaving Google reliant on Flash-tier models. On August 5, Demis Hassabis stepped down as DeepMind CEO to become its chairman and Alphabet’s chief scientist. Jeff Dean left to found Discovery Loop with Oriol Vinyals, Quoc Le and others. Koray Kavukcuoglu took over DeepMind’s day-to-day operations, reporting to Sundar Pichai.
Alphabet fell about 5% that day. Much of the commentary saw a shift toward selling compute and distributing AI applications. Tim O’Reilly drew a parallel with Westinghouse, suggesting Google might focus on infrastructure and wider AI adoption. Our view was different: Google would keep investing in frontier models, and a tighter focus would give it room to catch up. Argon’s results support that view. The training run began before the reorganization, however, so the effect of the management changes on R&D will only become clear over the next few model generations.
Early this year, some in the market considered a scenario in which Anthropic kept extending its lead in coding and the enterprise, the other labs gradually fell away, and the frontier ended up with a single player. Today, OpenAI and Anthropic still lead, Google is back in contention, and Meta and xAI continue to invest. Several labs remain in the race.
Google returns to frontier competition
Argon’s Artificial Analysis score is 23 points above Gemini 3.1 Pro and 12 above Gemini 3.8 Flash, previously Google’s highest-scoring model. It ranks first among 41 models on the Vals Index at 68.9%, versus roughly 67% for Sonnet 5.5 and Opus 5.5, and tops the LMArena text leaderboard at 1525.
Google disclosed on July 22 that its “most ambitious pretraining run yet” was in progress and announced Argon roughly ten weeks later. Our checks indicate the run used TPUs. We had pushed back on rumors of a move to Nvidia’s GB-series chips: Google has spent years optimizing its training systems around TPUs, and switching to GPUs would require extensive adaptation and revalidation. A training restart alone does not establish that there is a problem with the chips.
Pretraining scaling continues to deliver
Google has not disclosed the model’s size beyond calling this its “most ambitious pretraining run yet” and the model “significantly larger.” Our industry checks suggest a larger MoE base with a looped architecture: roughly 5–6 trillion total parameters, 400–500 billion active per token, and three loops. We estimate its effective scale at five to six times that of Gemini 3 Pro. This increases both model capacity and compute per token.
Many in the industry argued over the past year that pretraining scaling had run its course and that the focus should shift to post-training and inference. Gemini 4 expands both pretraining and inference compute. The output limit rises from 64,000 tokens to 1 million, and Long Decode Continuation lets a response and resume across calls. The model can spend more tokens working through long tasks, but completing them reliably still depends on tools and state management.
Long-context reasoning has improved as well. On GraphWalks with inputs of 256K–1M tokens, Argon scores 84.2%, versus 71.8% for Astra and 66.8% for Opus 5.5. This measures multi-hop reasoning over long inputs, a separate capability from generating longer outputs.
Argon uses markedly more tokens per task. Across the Artificial Analysis index, Argon generated an average of 62,000 output tokens per task, more than twice Astra’s 27,000. Long agent runs can also be expensive: a CUA-bench run on Vals costs $193.78 at standard pricing with a 262K output cap.