OpenAI official logo (public domain, Wikimedia Commons) — CryptoBriefing brand treatment
The new model climbs to #6 on the Agent Arena leaderboard at roughly half the price of its predecessor, intensifying the AI pricing war with Anthropic.
OpenAI just made its most aggressive cost-performance play yet. GPT-6 Sol (Max), released on September 22, 2026, posted a +7.7% net improvement across more than 4,000 real-world agentic sessions on the Agent Arena leaderboard, while operating at roughly half the API cost of its predecessor. The model now sits at #6 on the leaderboard, up from the #8 spot held by GPT-5.6 Sol (xHigh), which managed only a +6.2% improvement.
The numbers behind Sol’s leap #
GPT-6 Sol’s API pricing lands at $2 per million input tokens and $10 per million output tokens. That represents a 50% reduction from GPT-5.6 rates, a pricing move that transforms the economics of running agentic workflows at scale.
On the Confirmed Success metric, which measures tasks definitively completed rather than partially attempted, GPT-6 Sol scored +11.4% and ranked #4 overall. That’s a stronger showing than its composite leaderboard position suggests, indicating the model is particularly reliable when tasks demand completion rather than approximation.
On the DeepSWE v1.1 benchmark, a widely tracked measure of software engineering capability at maximum effort, GPT-6 Sol posted a score of 68.8%. That figure trails the top results from Anthropic’s Claude Fable 5 configurations. But context matters here: OpenAI claims Sol achieves its DeepSWE performance at around 80% less cost than the highest-performing Claude counterparts.
AI, tech, and the markets they move—in one daily briefing.
Daily. Free. Join 34,000+ readers across crypto, finance, and policy.
GPT-6 family expands its roster #
Sol didn’t arrive alone. OpenAI launched GPT-6 Luna alongside it on the same day, with both models joining GPT-6 Astra, which had already been released as part of the broader GPT-6 family. GPT-6 Sol and Luna specifically target scalable, cost-effective applications, including agentic workflows, professional automation pipelines, and always-on AI infrastructure.
The pricing war heats up #
Anthropic’s Claude Opus 5 and Claude Fable 5 remain formidable, with top-tier DeepSWE scores that exceed Sol’s 68.8%. The AutomationBench results, where Sol reportedly completed tasks at lower cost than certain Claude Opus 5 configurations, underscore the strategic intent: OpenAI is targeting the cost-per-completed-task metric rather than absolute benchmark rankings.
Disclosure: This article was edited by Diego Almada Lopez. For more information on how we create and review content, see our