GPT-6 Astra Achieves SOTA on ARC-AGI
OpenAI's GPT-6 Astra achieved state-of-the-art results on the ARC-AGI benchmark, scoring 63% on ARC-AGI-3 and 99% via a new provider adapter harness, surpassing human performance on 96% of ARC-AGI-3 l…
OpenAI's GPT-6 Astra achieved state-of-the-art results on the ARC-AGI benchmark, scoring 63% on ARC-AGI-3 and 99% via a new provider adapter harness, surpassing human performance on 96% of ARC-AGI-3 l…
Samsung's Tiny Recursive Model (TRM), with only 7 million parameters and two layers, outperforms frontier models like GPT-4 on the ARC-AGI-2 benchmark, suggesting that architecture and recursive logic…
Anthropic released Claude Fable 5.1 on September 1, achieving 90% coverage on the ARC-AGI-2 benchmark at 32% lower cost per task than its predecessor, Fable 5, which cost $5.45 per task. The model als…
Anthropic's Claude Fable 5.1 scores 97.5% on ARC-AGI-1 Semi-Private at $1.40 per task and 90.0% on ARC-AGI-2 Semi-Private at $4.49 per task at max effort, according to results published by Anthropic o…
Researchers introduced Meta^n, a recursive self-improvement system that keeps its meta-operation fixed and recurses on its input, achieving superior performance over prior self-improving agents on all…
Google's Gemini 3.7 Flash scored 84.6% on ARC-AGI-2 at $0.25 per task and 95.5% on ARC-AGI-1 at $0.12 per task, placing it within a few points of Claude Fable 5 (Max) and GPT-5.6 Sol (Max) at a fracti…
Google's Gemini 3.7 Flash scored 95.5% on ARC-AGI-1 Semi-Private at $0.12 per task and 84.6% on ARC-AGI-2 Semi-Private at $0.25 per task at high effort, according to results published by the ARC Prize…
The 2026 ARC-AGI Prize, now part of Vanguard AI Research, will award $1 million to the first team that achieves an 85% accuracy score on the ARC-AGI-2 benchmark, a test designed to measure fluid intel…
Google has announced Gemini 3.1 Pro, a new preview model designed for complex tasks with improved reasoning, rolling out across consumer, developer, and enterprise products. The model is available thr…
Researchers introduced PoTRE (Poly-Topological Reasoning Ensembles), a heterogeneous framework that decouples inference into four agents to improve complex reasoning in large language models. PoTRE ac…
A project called ARC-AGI-2 has achieved a 20.7% score using zero neural networks, focusing on LLM-free symbolic reasoning. The initiative also includes innovative iOS games controlled by mouth movemen…
Researchers present ARCANA, a collaborative multi-agent framework for solving ARC-AGI-2 tasks under strict test-time and hardware constraints. The framework decomposes tasks into iterative perception,…
Researchers introduced Regressive Plasticity Schedule (RPS), a two-stage post-training schedule that couples a learning-rate drop with a curriculum boundary between easier and harder data. Testing on …
A researcher built a model with a single parameter that achieved 100% accuracy on the ARC-AGI-2 benchmark, a million-dollar challenge designed to test reasoning models. The model used chaos theory and…