{"slug": "benchmarks-disagree-on-gpt-6-astra-but-its-human-beating-efficiency-on-arc-agi-3", "title": "Benchmarks disagree on GPT-6 Astra, but its human-beating efficiency on ARC-AGI-3 pulls Chollet’s AGI forecast forward", "summary": "OpenAI's GPT-6 Astra has surpassed average human efficiency on the ARC-AGI-3 benchmark for the first time, prompting ARC Prize chief François Chollet to move up his AGI forecast, though he stops short of calling it proof of AGI. Benchmark results are mixed: Epoch AI scores Astra at 169 points, while Artificial Analysis rates it no better than its predecessor and behind Claude Fable 5.1. Chollet said progress is running \"twice as fast\" as he expected.", "body_md": "OpenAI's GPT-6 Astra is drawing contradictory benchmark verdicts. Epoch AI puts it out in front with 169 points, while Artificial Analysis rates it no better than its predecessor and behind Claude Fable 5.1. The biggest surprise comes from ARC-AGI-3, where Astra works more efficiently than the average human for the first time. ARC Prize chief François Chollet doesn't call this proof of AGI, but he does see the progress running \"twice as fast\" as he expected, and he's moving up his AGI forecast.\n\nThe article [Benchmarks disagree on GPT-6 Astra, but its human-beating efficiency on ARC-AGI-3 pulls Chollet’s AGI forecast forward](https://the-decoder.com/benchmarks-disagree-on-gpt-6-astra-but-its-human-beating-efficiency-on-arc-agi-3-pulls-chollets-agi-forecast-forward/) appeared first on [The Decoder](https://the-decoder.com).", "url": "https://wpnews.pro/news/benchmarks-disagree-on-gpt-6-astra-but-its-human-beating-efficiency-on-arc-agi-3", "canonical_source": "https://the-decoder.com/benchmarks-disagree-on-gpt-6-astra-but-its-human-beating-efficiency-on-arc-agi-3-pulls-chollets-agi-forecast-forward/", "published_at": "2026-09-04 11:07:36+00:00", "updated_at": "2026-09-04 11:23:23.980623+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-research", "ai-products"], "entities": ["OpenAI", "GPT-6 Astra", "Epoch AI", "Artificial Analysis", "Claude Fable 5.1", "ARC-AGI-3", "François Chollet", "ARC Prize"], "alternates": {"html": "https://wpnews.pro/news/benchmarks-disagree-on-gpt-6-astra-but-its-human-beating-efficiency-on-arc-agi-3", "markdown": "https://wpnews.pro/news/benchmarks-disagree-on-gpt-6-astra-but-its-human-beating-efficiency-on-arc-agi-3.md", "text": "https://wpnews.pro/news/benchmarks-disagree-on-gpt-6-astra-but-its-human-beating-efficiency-on-arc-agi-3.txt", "jsonld": "https://wpnews.pro/news/benchmarks-disagree-on-gpt-6-astra-but-its-human-beating-efficiency-on-arc-agi-3.jsonld"}}