{"slug": "peter-steinberger-called-out-anthropic-for-posting-arc-agi-3-results-that-openai", "title": "Peter Steinberger called out Anthropic for posting ARC-AGI-3 results that OpenAI quickly made look absurd", "summary": "Anthropic's Claude Opus 5 scored 30.16% on the ARC-AGI-3 Public Demo set, according to ARC Prize results covered by TechTimes and Digital Applied, clearing five environments no AI model had previously solved and surpassing the previous top score of 7.78% by OpenAI's GPT-5.6 Sol. However, claims that OpenAI then boosted GPT-5.6 Sol to 38.3% by enabling retained reasoning and compaction could not be verified, as the original OpenAI post was not found in live search, and a quote attributed to Peter Steinberger, an OpenAI employee and creator of OpenClaw, did not surface either.", "body_md": "*Claude Opus 5's ARC-AGI-3 score is real. The clean story about OpenAI instantly beating it with two API settings is not solid enough to print as fact.*\n\nThe benchmark fight is useful, but only if you keep the ground under your feet. ARC Prize's July 24, 2026 result for Claude Opus 5 has a clear public trail: 30.16%, usually rounded to 30.2%, on ARC-AGI-3. That part checks out. The claim that OpenAI then jumped GPT-5.6 Sol from 7.8% to 38.3% with retained reasoning and compaction is a different kind of claim. You need the original OpenAI post for that. A live search did not surface it.\n\nThat is not a small problem. It is the whole problem.\n\nAccording to ARC Prize results covered by TechTimes and Digital Applied, Claude Opus 5 scored 30.16% on the ARC-AGI-3 Public Demo set, clearing five environments no AI model had previously solved. The same reporting says the previous top score was GPT-5.6 Sol at roughly 7.8%. OpenAI's own GPT-5.6 launch page also lists GPT-5.6 Sol at 7.78% on ARC-AGI-3. Those are the numbers you can stand behind.\n\nThe original draft went further. It attributed a sharp quote to Peter Steinberger, known online as @steipete, and said he asked whether anyone at Anthropic had stopped to wonder why the numbers looked absurd before posting a victory tweet. Steinberger is real. TechCrunch identified him in April as OpenClaw's creator and an OpenAI employee, and his public GitHub profile connects him to OpenClaw and OpenAI. But the exact quote in the draft did not show up in live search, including searches for the quoted language, his handle, Anthropic, and ARC-AGI-3.\n\nYou don't print that quote without the post.\n\nThe same rule applies to the supposed OpenAI rebuttal titled \"How enabling two settings tripled our scores on the ARC-AGI-3 benchmark.\" If that post exists behind a page search could not reach, it still needs to be in the brief or linked in the story. Without it, the article cannot state as fact that OpenAI enabled retained reasoning and compaction, ran GPT-5.6 Sol through its Responses API, and reached 38.3% while using six times fewer output tokens. Those details are too specific to carry on vibes.\n\n## The verified result still matters\n\nClaude Opus 5's verified score is enough of a story on its own. ARC-AGI-3 was launched as an interactive benchmark where humans solve the environments and frontier models initially scored below 1%. A move from the earlier GPT-5.6 Sol figure of 7.78% to Opus 5 at 30.16% is a genuine leap under ARC Prize's public scoring process. You don't need a fake clean controversy to make that interesting.\n\nThe result also has useful texture. TechTimes reported that Opus 5 solved five previously unsolved public demo environments and produced an explicit reflection equation during one run. That is the kind of detail readers can use. It tells you more than a launch-chart boast because it points to behavior inside the test, not only the number printed at the end.\n\nBut benchmark stories punish carelessness. A vendor-run result, a third-party verified result, a public demo set, a private holdout, a high-effort run and a max-effort run can all look like one leaderboard if you blur them together. They are not the same thing. If you're choosing models, buying enterprise AI, or just trying to understand whether a lab has really moved the field, you have to read the harness.\n\nFrankly, this is where the draft got too eager. It had the right instinct, skepticism toward a victory lap built on benchmark optics, but it used unverified attribution to make the point sharper than the facts allowed. The safer, stronger article is simpler: Anthropic's Opus 5 result was independently verified at about 30.2%; OpenAI's public GPT-5.6 page still shows 7.78% on ARC-AGI-3; any claim that a settings change pushed that to 38.3% needs a direct source before StartupFortune treats it as fact.\n\nThat leaves a less dramatic story, but a better one. ARC-AGI-3 is now important enough that every scoring condition matters, and every lab has an incentive to make its number look decisive. Your job as a reader is not to cheer the biggest bar on the chart. Read the small print first.\n\n**Also read:** [Xscape Photonics raised $81 million to replace copper wiring with light inside AI data centers](https://startupfortune.com/xscape-photonics-raised-81-million-to-replace-copper-wiring-with-light-inside-ai-data-centers/) • [Unitree Robotics heads into its Shanghai IPO with a $619 million target and a US ban hanging over it](https://startupfortune.com/unitree-robotics-heads-into-its-shanghai-ipo-with-a-619-million-target-and-a-us-ban-hanging-over-it/) • [Simile raises $200 million at a $2 billion valuation to replace focus groups with AI simulated humans](https://startupfortune.com/simile-raises-200-million-at-a-2-billion-valuation-to-replace-focus-groups-with-ai-simulated-humans/)", "url": "https://wpnews.pro/news/peter-steinberger-called-out-anthropic-for-posting-arc-agi-3-results-that-openai", "canonical_source": "https://startupfortune.com/peter-steinberger-called-out-anthropic-for-posting-arc-agi-3-results-that-openai-quickly-made-look-absurd/", "published_at": "2026-07-31 03:55:13+00:00", "updated_at": "2026-07-31 04:06:27.373572+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-research", "ai-policy"], "entities": ["Anthropic", "Claude Opus 5", "OpenAI", "GPT-5.6 Sol", "ARC Prize", "ARC-AGI-3", "Peter Steinberger", "TechTimes"], "alternates": {"html": "https://wpnews.pro/news/peter-steinberger-called-out-anthropic-for-posting-arc-agi-3-results-that-openai", "markdown": "https://wpnews.pro/news/peter-steinberger-called-out-anthropic-for-posting-arc-agi-3-results-that-openai.md", "text": "https://wpnews.pro/news/peter-steinberger-called-out-anthropic-for-posting-arc-agi-3-results-that-openai.txt", "jsonld": "https://wpnews.pro/news/peter-steinberger-called-out-anthropic-for-posting-arc-agi-3-results-that-openai.jsonld"}}