{"slug": "gemini-4-argon-with-antigravity-places-ahead-of-gpt-6-1-sol-on-codex-on-analysis", "title": "Gemini 4 Argon With Antigravity Places Ahead Of GPT 6.1 Sol On Codex On Artificial Analysis Coding Agent Index", "summary": "Google's Gemini 4 Argon, run in Google's Antigravity CLI, scored 64 on the Artificial Analysis Coding Agent Index (v1.5), one point ahead of OpenAI's GPT-6.1 Sol (xhigh) running in Codex at 63, according to Artificial Analysis. Argon placed third overall behind Anthropic's Sonnet 5.5 and Opus 5.5 in Claude Code, and ahead of OpenAI's flagship GPT-6 Astra at 62 in Codex at max effort. Argon is not yet publicly available, having been rolled out only to select users, and Artificial Analysis measured its cost at $1.99 per task on the Intelligence Index versus $0.72 for GPT-6.1 Sol (max).", "body_md": "Google might’ve finally solved its AI coding bugbear.\n\nGoogle’s new [Gemini 4 Argon](https://officechai.com/ai/gemini-4-argon-benchmarks/) has edged past OpenAI’s GPT-6.1 Sol on the Artificial Analysis Coding Agent Index, when each model is run in its maker’s own coding harness. Argon, running in Google’s Antigravity CLI, scores 64 on the index, one point ahead of GPT-6.1 Sol (xhigh) running in Codex, which scores 63.\n\nThe result is interesting because it benchmarks combinations of models and agent harnesses, reflecting how developers actually use AI to write code. Google’s tool is now in the same conversation as Codex and Claude Code, a place Gemini-based coding tools haven’t often been.\n\n## How The Index Shakes Out\n\nThe Artificial Analysis Coding Agent Index (v1.5) is an equal-weight average of three benchmarks: DeepSWE v1.1, Terminal-Bench 4.0 and SWE-Atlas-QnA. Argon with Antigravity lands third, behind only Anthropic’s Sonnet 5.5 and Opus 5.5 in Claude Code. It sits ahead of both OpenAI entries on the chart, including [GPT-6 Astra](https://officechai.com/ai/openai-gpt-6-astra-scores-a-disappointing-61-on-artificial-analysis-intelligence-index-same-as-gpt-5-6-sol/), OpenAI’s flagship, which scores 62 in Codex at max effort. Argon also edges out Claude Fable 5.1 in Claude Code with fallback, at 62.\n\nThe caveat is that Argon isn’t yet publicly available — Google has so far rolled it out only to select users, with a broader release to come, so developers can’t yet reproduce the result on their own.\n\n## GPT-6.1 Sol’s Coding Agent Showing\n\n[GPT-6.1 Sol](https://officechai.com/ai/gpt-6-1-sol/) was released on September 29 as an upgrade to GPT-6 Sol, with OpenAI saying it nearly matches Astra’s intelligence on agentic coding, computer use and professional work at a fifth of the price. The Coding Agent Index result lines up with that claim: at 63, Sol in Codex sits a point above Astra in the same harness.\n\nThe Argon-over-Sol ordering is a narrow one, though. A single point on a composite index is a thin margin, and the two are effectively in the same tier.\n\n## Argon’s Strengths And Gaps\n\nArtificial Analysis’ [separate evaluation of Argon](https://officechai.com/ai/gemini-4-argon-ties-with-gpt-6-astra-on-artificial-analysis-intelligence-index-at-a-40-cheaper-price/) helps explain the result. On DeepSWE, one of the three components of the coding agent index, Argon scored 77.5% in the firm’s independent testing, ahead of Claude Sonnet 5.5 (max) at 71.3%, Claude Opus 5.5 at 69.5%, GPT-6 Astra at 68.5% and Claude Fable 5.1 at 59.4%.\n\nTerminal-based work is less flattering. On Terminal-Bench 4.0, Argon scores 57.1%, which is a big leap from the 4.0% of Gemini 3.1 Pro Preview, but still behind Claude Sonnet 5.5 (63.6%), Claude Opus 5.5 (59.6%) and GPT-6 Astra (59.1%). It is ahead of GPT-6.1 Sol at 56.1% and Fable 5.1 at 52.0%. That fits the picture from the [broader Intelligence Index results](https://officechai.com/ai/gemini-is-back-to-the-frontier-with-gemini-4-argon/), where Argon scores 53, a point above GPT-6.1 Sol at 52, while still trailing several rivals on terminal-based coding.\n\n## The Cost Question\n\nArgon’s lead isn’t free. On Artificial Analysis’ cost-per-task chart for the Intelligence Index, Gemini 4 Argon (high) comes in at $1.99 per task, compared with $0.72 for GPT-6.1 Sol (max). That’s roughly 2.7 times as much, even at Argon’s discounted introductory rate of $2 per million input tokens and $10 per million output tokens. Google has said those rates will eventually double.\n\nAnthropic’s models sit at the other end of the scale on that chart. Claude Opus 5.5 (max, with fallback) costs $5.98 per task, Claude Fable 5.1 $7.63, and Claude Sonnet 5.5 $7.67. So while Anthropic’s agents lead the coding index, Argon delivers a score within a few points of the top at a fraction of their per-task cost on the Intelligence Index. Note that the cost chart measures the Intelligence Index, not the Coding Agent Index, so it is an indication of relative pricing rather than a direct coding cost comparison.\n\n## What It Means\n\nFor Google, which spent months being questioned over its position on the leaderboards, this is another data point that Gemini is back near the frontier. It also validates the company’s push to consolidate its developer tooling around Antigravity, which replaced Gemini CLI as its terminal agent earlier this year.\n\nThe coding agent space is increasingly a race between harness-and-model pairings rather than models alone, and the top of this chart is crowded. Anthropic still leads, but Google’s first Gemini 4 model has shown it can land between Anthropic and OpenAI. The real test will come when Argon is available to everyone, and developers can see whether the numbers hold up in day-to-day use. Earlier this year, [Cursor’s Composer 2.5](https://officechai.com/ai/cursors-composer-2-5-places-3rd-in-artificial-analysis-coding-agent-index-is-10-60x-cheaper-than-variants-above-it/) showed that this index can also reward cheaper agents, so cost-per-task could be as important as score as the field matures.", "url": "https://wpnews.pro/news/gemini-4-argon-with-antigravity-places-ahead-of-gpt-6-1-sol-on-codex-on-analysis", "canonical_source": "https://officechai.com/ai/gemini-4-argon-with-antigravity-places-ahead-of-gpt-6-1-sol-on-codex-on-artificial-analysis-coding-agent-index/", "published_at": "2026-10-05 10:39:26+00:00", "updated_at": "2026-10-05 11:16:57.760072+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-agents", "ai-tools", "ai-products"], "entities": ["Google", "Gemini 4 Argon", "Antigravity", "OpenAI", "GPT-6.1 Sol", "Codex", "Artificial Analysis", "Anthropic"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/gemini-4-argon-with-antigravity-places-ahead-of-gpt-6-1-sol-on-codex-on-analysis", "markdown": "https://wpnews.pro/news/gemini-4-argon-with-antigravity-places-ahead-of-gpt-6-1-sol-on-codex-on-analysis.md", "text": "https://wpnews.pro/news/gemini-4-argon-with-antigravity-places-ahead-of-gpt-6-1-sol-on-codex-on-analysis.txt", "jsonld": "https://wpnews.pro/news/gemini-4-argon-with-antigravity-places-ahead-of-gpt-6-1-sol-on-codex-on-analysis.jsonld"}}