{"slug": "gemini-4-argon-tops-the-benchmarks-you-can-t-use-it-here-s-what-actually-matters", "title": "Gemini 4 Argon tops the benchmarks. You can't use it. Here's what actually matters.", "summary": "Google released Gemini 4 Argon on September 30, a frontier model pitched at long-running software engineering, enterprise knowledge work and cyber defense, but it is not generally available as Google iterates on guardrails and coordinates with the US government's voluntary pre-release access process. Independent evaluator Vals ranks Argon first on its index at 68.9% and roughly $15.68 per test versus $32.14 for Claude Opus 5.5, while the model loses to Opus 5.5 by 9 points on Terminal-Bench 4.0 and to GPT-6 Astra by 10 points on FrontierSWE v2. Introductory pricing is $2/$10 per million input/output tokens with cached input 95% off, though the article notes the pricing promise expires and the model remains uncallable for most developers.", "body_md": "Google dropped Gemini 4 Argon on September 30 and Hacker News lost its mind: 1,300+ points, 870+ comments, the top of the front page.\n\nThen everyone tried to use it and hit a wall. Argon is **not** generally available. Let's separate the signal from the launch-day noise.\n\nArgon is Google's new frontier model, pitched at long-running workflows: software engineering, enterprise knowledge work (law, finance), and cyber defense.\n\nAccess today:\n\nGoogle says it is iterating on guardrails first and coordinating with the US government's voluntary pre-release access process. Translation: the cyber capability is the reason it's gated. Wiz reportedly used it to find a critical healthcare-software vulnerability that earlier frontier models missed.\n\n| Benchmark | Argon | GPT-6 Astra | Claude Opus 5.5 | \n|---|---|---|---|\n| DeepSWE v1.1 | **77.9%** | 74.1% | 74.2% | \n| Terminal-Bench 4.0 | 57.4% | n/a | **66.4%** | \n| FrontierSWE v2 | 55.0% | **65.5%** | n/a | \n| Vals Index | **68.9%** | n/a | 67.0% | \n| GraphWalks (256K-1M) | **84.2%** | 71.8% | 66.8% | \n| OSWorld-2.0 | 69.2% | **72.6%** | n/a | \n\nIndependent Vals puts Argon #1 on its index (68.9%) at about $15.68 per test, versus $32.14 for Opus 5.5.\n\nRead that table carefully. Argon wins on repo-level SWE tasks and long-context retrieval. It **loses** on terminal-driven agent work (Terminal-Bench 4.0, by 9 points to Opus) and on FrontierSWE v2 (10 points to Astra). \"Beats everyone on most benchmarks\" is true. \"Best coding model\" is not.\n\nIntroductory pricing is **$2 / $10 per million tokens** (input/output), with cached input 95% off. Reported comparisons:\n\nAt intro pricing, that's 5x cheaper than Astra with a higher score on DeepSWE. Even at the doubled price it matches Opus. If it holds up outside Google's evals, this is a margin story for anyone running agents at volume: the cost per resolved task is what your CFO sees, not the leaderboard.\n\nThe thread is less about Argon and more about the **harness**:\n\nPeter Yang's take is the cleanest summary: Google cooked on the model, now it needs to compete on the coding harness and the personal-agent product.\n\nModel quality is converging. What differentiates now is three boring things:\n\nArgon scores well on 1, is unproven on 2, and currently fails 3 for almost everyone.\n\nDon't rewrite anything. Do this instead:\n\n```\n# Keep your model behind one seam so swapping is a config change\nMODELS = {\n    \"default\": \"claude-opus-5-5\",\n    \"cheap_bulk\": \"gemini-4-argon\",   # flip on when GA\n}\n\ndef pick(task):\n    return MODELS[\"cheap_bulk\"] if task.is_bulk_swe else MODELS[\"default\"]\n```\n\nArgon might be the best price-to-performance frontier model on the board. It's also a model you can't call yet, in a harness developers don't like, with a pricing promise that expires. Respect the benchmark, ignore the hype, and keep your abstraction layer clean.\n\n*Sources: Hacker News discussion, Vals.ai, OfficeChai, Droid Life, tbreak. Figures reported September 30, 2026; Google's benchmarks are unverified by independent testers.*", "url": "https://wpnews.pro/news/gemini-4-argon-tops-the-benchmarks-you-can-t-use-it-here-s-what-actually-matters", "canonical_source": "https://dev.to/ashraf_chowdury09/gemini-4-argon-tops-the-benchmarks-you-cant-use-it-heres-what-actually-matters-436f", "published_at": "2026-10-01 09:02:31+00:00", "updated_at": "2026-10-01 09:14:20.021822+00:00", "lang": "en", "topics": ["large-language-models", "ai-products", "ai-agents", "ai-tools"], "entities": ["Google", "Gemini 4 Argon", "GPT-6 Astra", "Claude Opus 5.5", "Vals", "Wiz", "Hacker News", "Peter Yang"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/gemini-4-argon-tops-the-benchmarks-you-can-t-use-it-here-s-what-actually-matters", "markdown": "https://wpnews.pro/news/gemini-4-argon-tops-the-benchmarks-you-can-t-use-it-here-s-what-actually-matters.md", "text": "https://wpnews.pro/news/gemini-4-argon-tops-the-benchmarks-you-can-t-use-it-here-s-what-actually-matters.txt", "jsonld": "https://wpnews.pro/news/gemini-4-argon-tops-the-benchmarks-you-can-t-use-it-here-s-what-actually-matters.jsonld"}}