Google dropped Gemini 4 Argon on September 30 and Hacker News lost its mind: 1,300+ points, 870+ comments, the top of the front page.
Then everyone tried to use it and hit a wall. Argon is not generally available. Let's separate the signal from the launch-day noise.
Argon is Google's new frontier model, pitched at long-running workflows: software engineering, enterprise knowledge work (law, finance), and cyber defense.
Access today:
Google says it is iterating on guardrails first and coordinating with the US government's voluntary pre-release access process. Translation: the cyber capability is the reason it's gated. Wiz reportedly used it to find a critical healthcare-software vulnerability that earlier frontier models missed.
| Benchmark | Argon | GPT-6 Astra | Claude Opus 5.5 |
|---|---|---|---|
| DeepSWE v1.1 | 77.9% | 74.1% | 74.2% |
| Terminal-Bench 4.0 | 57.4% | n/a | 66.4% |
| FrontierSWE v2 | 55.0% | 65.5% | n/a |
| Vals Index | 68.9% | n/a | 67.0% |
| GraphWalks (256K-1M) | 84.2% | 71.8% | 66.8% |
| OSWorld-2.0 | 69.2% | 72.6% | n/a |
Independent Vals puts Argon #1 on its index (68.9%) at about $15.68 per test, versus $32.14 for Opus 5.5.
Read that table carefully. Argon wins on repo-level SWE tasks and long-context retrieval. It loses on terminal-driven agent work (Terminal-Bench 4.0, by 9 points to Opus) and on FrontierSWE v2 (10 points to Astra). "Beats everyone on most benchmarks" is true. "Best coding model" is not.
Introductory pricing is $2 / $10 per million tokens (input/output), with cached input 95% off. Reported comparisons:
At intro pricing, that's 5x cheaper than Astra with a higher score on DeepSWE. Even at the doubled price it matches Opus. If it holds up outside Google's evals, this is a margin story for anyone running agents at volume: the cost per resolved task is what your CFO sees, not the leaderboard.
The thread is less about Argon and more about the harness:
Peter Yang's take is the cleanest summary: Google cooked on the model, now it needs to compete on the coding harness and the personal-agent product.
Model quality is converging. What differentiates now is three boring things:
Argon scores well on 1, is unproven on 2, and currently fails 3 for almost everyone.
Don't rewrite anything. Do this instead:
MODELS = {
"default": "claude-opus-5-5",
"cheap_bulk": "gemini-4-argon", # flip on when GA
}
def pick(task):
return MODELS["cheap_bulk"] if task.is_bulk_swe else MODELS["default"]
Argon might be the best price-to-performance frontier model on the board. It's also a model you can't call yet, in a harness developers don't like, with a pricing promise that expires. Respect the benchmark, ignore the hype, and keep your abstraction layer clean.
Sources: Hacker News discussion, Vals.ai, OfficeChai, Droid Life, tbreak. Figures reported September 30, 2026; Google's benchmarks are unverified by independent testers.