{"slug": "do-coding-agents-need-executable-world-models-simplification-and-verification-to", "title": "Do Coding Agents Need Executable World Models, Simplification, and Verification to Solve ARC-AGI-3?", "summary": "A new study from arXiv (2607.15439v1) finds that a verification-based coding agent using gpt-5.6-sol fully solves every public ARC-AGI-3 game at both high and max reasoning effort, achieving about 99% RHAE with fewer than half the total actions of the human baseline. The research, which tested four nested Codex-based agents across multiple GPT models, shows that the complete verification treatment (executable world modeling, scheduled simplification, and exact replay verification) ranks first in all settings, though it uses substantially more resources. The authors caution that because the model postdates these games and held-out performance remains untested, this result represents saturation of the public set only.", "body_md": "arXiv:2607.15439v1 Announce Type: new\nAbstract: Our previous ARC-AGI-3 agent bundled executable world modeling, scheduled simplification, and exact replay verification, leaving unclear which idea accounted for its performance. We address this attribution question with four nested Codex-based agents: a textual baseline; a flexible-interface executable world model without replay verification; the same executable model with scheduled simplification; and a fixed-interface verification treatment that retains simplification and requires exact reproduction of recorded observations. The main study evaluates all four agents with gpt-5.4 and gpt-5.5 at high and xhigh reasoning effort on the public ARC-AGI-3 games. Exploratory follow-ups evaluate the textual and verification variants with gpt-5.6-sol at xhigh and max. The most robust result is that every agent variant improves with a stronger model and with greater reasoning effort. Within each model-effort setting, differences among variants are smaller than anticipated, while the effects of individual components vary across settings. Requiring a persistent executable deliverable is not universally beneficial: the textual variant outperforms the flexible-interface executable variant in both gpt-5.5 settings. Simplification improves performance in three of the four model-effort settings, with the weakest setting as the only exception. The complete verification treatment ranks first in all four settings, although it uses substantially more resources. In the gpt-5.6-sol follow-up, the verification variant fully solves every public game at both reasoning efforts, achieves about 99% RHAE, and uses fewer than half the total actions of the human baseline. Because the model postdates these games and held-out performance remains untested, this result should be interpreted as saturation of the public set only.", "url": "https://wpnews.pro/news/do-coding-agents-need-executable-world-models-simplification-and-verification-to", "canonical_source": "https://arxiv.org/abs/2607.15439", "published_at": "2026-07-20 04:00:00+00:00", "updated_at": "2026-07-20 13:52:45.492437+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "ai-agents", "ai-research"], "entities": ["arXiv", "Codex", "ARC-AGI-3", "gpt-5.4", "gpt-5.5", "gpt-5.6-sol"], "alternates": {"html": "https://wpnews.pro/news/do-coding-agents-need-executable-world-models-simplification-and-verification-to", "markdown": "https://wpnews.pro/news/do-coding-agents-need-executable-world-models-simplification-and-verification-to.md", "text": "https://wpnews.pro/news/do-coding-agents-need-executable-world-models-simplification-and-verification-to.txt", "jsonld": "https://wpnews.pro/news/do-coding-agents-need-executable-world-models-simplification-and-verification-to.jsonld"}}