{"slug": "claims-of-gpt-6-astra-scoring-98-6-on-arc-agi-3-dont-hold-up-to-scrutiny", "title": "Claims of GPT-6 Astra scoring 98.6% on ARC-AGI-3 don’t hold up to scrutiny", "summary": "A viral claim that OpenAI's GPT-6 Astra scored 98.6% on the ARC-AGI-3 benchmark is unsupported by verified evidence, according to Crypto Briefing. The current top score on ARC-AGI-3 is Anthropic's Claude Opus 5 at roughly 30.2%, while OpenAI's GPT-5.6 Sol scored 7.78% under official testing and 38.3% with custom Responses API settings. OpenAI has not confirmed a release date or branding for GPT-6 as of early September 2026.", "body_md": "Photo: Tima Miroshnichenko / Pexels\n\n# Claims of GPT-6 Astra scoring 98.6% on ARC-AGI-3 don’t hold up to scrutiny\n\nThe top model on the ARC-AGI-3 leaderboard currently scores around 30%, making the viral 98.6% claim look wildly disconnected from reality\n\nA claim circulating on social media that OpenAI’s GPT-6 Astra scored 98.6% on the ARC-AGI-3 benchmark would, if true, represent one of the most significant leaps in AI capability ever recorded. The problem: there’s no verified evidence to back it up.\n\nThe ARC-AGI-3 benchmark, which launched on March 25, 2026, is specifically designed to test whether AI models can navigate interactive environments without instructions or predefined objectives. When models first encountered the benchmark, scores came in below 1%.\n\n## What the leaderboard actually shows\n\nThe current top performer on ARC-AGI-3 is Anthropic’s Claude Opus 5, sitting at roughly 30.2%.\n\nOpenAI’s own GPT-5.6 Sol, its most recent model with verified benchmark results, managed 7.78% under official testing conditions. When OpenAI used its custom Responses API settings, that number climbed to 38.3%.\n\nNo confirmed ARC-AGI-3 scores exist for a model called Astra. As of early September 2026, OpenAI hasn’t even established a confirmed release date or official branding for GPT-6.\n\n## What we actually know about Astra\n\nOpenAI previewed Astra around August 1, 2026. The model earned recognition for solving approximately 10 open math problems, which are unsolved challenges that have stumped mathematicians.\n\nThe gap between Astra’s demonstrated capabilities and the claimed 98.6% score suggests one of two scenarios. Either the claim conflates or inflates GPT-5.6 Sol’s results, or it originates from an unverified source making assertions about Astra’s performance without supporting data.\n\n## Why ARC-AGI-3 matters\n\nThe ARC Prize Foundation designed ARC-AGI-3 as a successor to earlier versions of the benchmark, each iteration becoming harder to game through memorization or pattern-matching. The benchmark uses a scoring methodology called Relative Human Action Efficiency (RHAE), which allows for a comparative evaluation of AI models against human baselines in interactive settings.\n\nWhen initial models scored below 1%, it underscored just how difficult this benchmark is. The progression to Claude Opus 5’s 30.2% over roughly six months represents meaningful but incremental progress.\n\n## The competitive landscape and market implications\n\nAnthropic’s Claude Opus 5 holds the current ARC-AGI-3 lead, while OpenAI’s GPT-5.6 Sol shows competitive performance under optimized conditions.\n\nAstra’s mathematical reasoning achievements are noteworthy, and OpenAI’s standing in ARC-AGI-3 specifically, based on available data, places it behind Anthropic’s leading model.\n\n**Disclosure:** This article was edited by Editorial Team. For more information on how we create and review content, see our\n\n[Editorial Policy](https://cryptobriefing.com/editorial-policy/).", "url": "https://wpnews.pro/news/claims-of-gpt-6-astra-scoring-98-6-on-arc-agi-3-dont-hold-up-to-scrutiny", "canonical_source": "https://cryptobriefing.com/gpt-6-astra-arc-agi-3-claims-unverified/", "published_at": "2026-09-03 18:39:15+00:00", "updated_at": "2026-09-03 18:55:46.920643+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-research", "large-language-models"], "entities": ["OpenAI", "GPT-6 Astra", "ARC-AGI-3", "Anthropic", "Claude Opus 5", "GPT-5.6 Sol", "ARC Prize Foundation"], "alternates": {"html": "https://wpnews.pro/news/claims-of-gpt-6-astra-scoring-98-6-on-arc-agi-3-dont-hold-up-to-scrutiny", "markdown": "https://wpnews.pro/news/claims-of-gpt-6-astra-scoring-98-6-on-arc-agi-3-dont-hold-up-to-scrutiny.md", "text": "https://wpnews.pro/news/claims-of-gpt-6-astra-scoring-98-6-on-arc-agi-3-dont-hold-up-to-scrutiny.txt", "jsonld": "https://wpnews.pro/news/claims-of-gpt-6-astra-scoring-98-6-on-arc-agi-3-dont-hold-up-to-scrutiny.jsonld"}}