Thinking with Looped Flows
Researchers Ayhan Suleymanzade and co-authors submitted a paper to arXiv on 10 September 2026 proposing "looped flows," a training method that uses local denoising objectives with progressively decrea…
Researchers Ayhan Suleymanzade and co-authors submitted a paper to arXiv on 10 September 2026 proposing "looped flows," a training method that uses local denoising objectives with progressively decrea…
OpenAI's GPT-6 Astra achieved a 99.95% score on the ARC-AGI-3 Semi-Private benchmark using the Provider Adapter harness at high reasoning, with a cost of $18,817, according to results published by Ope…
Samsung's Tiny Recursive Model (TRM), with only 7 million parameters and two layers, outperforms frontier models like GPT-4 on the ARC-AGI-2 benchmark, suggesting that architecture and recursive logic…
Anthropic's Claude Fable 5.1 scores 97.5% on ARC-AGI-1 Semi-Private at $1.40 per task and 90.0% on ARC-AGI-2 Semi-Private at $4.49 per task at max effort, according to results published by Anthropic o…
A developer trained a small transformer from scratch in 1.5 hours on an NVIDIA RTX 5090 for 67 cents of compute, scoring 44% on the ARC-AGI-1 benchmark and 7% on ARC-2, matching the performance of TRM…
CodeSOTA, an open data terminal for reinforcement learning environments and state-of-the-art models, has launched a registry tracking 9,102 results, 163 models, 371 datasets, and 9 capability areas, w…
Researchers introduced BDH-CQ, a reasoning model combining in-context learning with recurrent latent reasoning, achieving 29.5% pass@2 on the ARC-AGI-1 evaluation set with a 150M-parameter configurati…
Google's Gemini 3.7 Flash scored 84.6% on ARC-AGI-2 at $0.25 per task and 95.5% on ARC-AGI-1 at $0.12 per task, placing it within a few points of Claude Fable 5 (Max) and GPT-5.6 Sol (Max) at a fracti…
Google's Gemini 3.7 Flash scored 95.5% on ARC-AGI-1 Semi-Private at $0.12 per task and 84.6% on ARC-AGI-2 Semi-Private at $0.25 per task at high effort, according to results published by the ARC Prize…
BDH-CQ, a 150M-parameter model, achieved a 29.5% pass@2 score on the ARC-AGI-1 benchmark at a cost of $0.00070 per task, demonstrating that recurrent latent state reasoning can outperform larger model…
A 150-million-parameter recurrent latent reasoning model has achieved a 29.5% score on the ARC-AGI-1 benchmark, a result that challenges the cost-to-accuracy frontier typically associated with much la…
Thinking Machines' 150M-parameter BDH-CQ reasoning model, which combines in-context learning with recurrent latent reasoning, achieves 29.5% pass@2 on the public ARC-AGI-1 evaluation set at a computed…
Pathway, a startup founded by Zuzanna Stamirowska, unveiled a 150-million parameter reasoning model, BDH-CQ, claiming it achieves comparable performance to leading frontier models at a fraction of the…
Researchers developed cost-effective agent harnesses for the ARC-AGI-1 abstract reasoning benchmark, achieving 67.25% pass@2 at $0.62 per task using an open-weight model without fine-tuning. The Refle…
Researchers introduced Regressive Plasticity Schedule (RPS), a two-stage post-training schedule that couples a learning-rate drop with a curriculum boundary between easier and harder data. Testing on …