cd /news/artificial-intelligence/do-coding-agents-need-executable-wor… · home topics artificial-intelligence article
[ARTICLE · art-65473] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Do Coding Agents Need Executable World Models, Simplification, and Verification to Solve ARC-AGI-3?

A new study from arXiv (2607.15439v1) finds that a verification-based coding agent using gpt-5.6-sol fully solves every public ARC-AGI-3 game at both high and max reasoning effort, achieving about 99% RHAE with fewer than half the total actions of the human baseline. The research, which tested four nested Codex-based agents across multiple GPT models, shows that the complete verification treatment (executable world modeling, scheduled simplification, and exact replay verification) ranks first in all settings, though it uses substantially more resources. The authors caution that because the model postdates these games and held-out performance remains untested, this result represents saturation of the public set only.

read1 min views5 publishedJul 20, 2026

arXiv:2607.15439v1 Announce Type: new Abstract: Our previous ARC-AGI-3 agent bundled executable world modeling, scheduled simplification, and exact replay verification, leaving unclear which idea accounted for its performance. We address this attribution question with four nested Codex-based agents: a textual baseline; a flexible-interface executable world model without replay verification; the same executable model with scheduled simplification; and a fixed-interface verification treatment that retains simplification and requires exact reproduction of recorded observations. The main study evaluates all four agents with gpt-5.4 and gpt-5.5 at high and xhigh reasoning effort on the public ARC-AGI-3 games. Exploratory follow-ups evaluate the textual and verification variants with gpt-5.6-sol at xhigh and max. The most robust result is that every agent variant improves with a stronger model and with greater reasoning effort. Within each model-effort setting, differences among variants are smaller than anticipated, while the effects of individual components vary across settings. Requiring a persistent executable deliverable is not universally beneficial: the textual variant outperforms the flexible-interface executable variant in both gpt-5.5 settings. Simplification improves performance in three of the four model-effort settings, with the weakest setting as the only exception. The complete verification treatment ranks first in all four settings, although it uses substantially more resources. In the gpt-5.6-sol follow-up, the verification variant fully solves every public game at both reasoning efforts, achieves about 99% RHAE, and uses fewer than half the total actions of the human baseline. Because the model postdates these games and held-out performance remains untested, this result should be interpreted as saturation of the public set only.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/do-coding-agents-nee…] indexed:0 read:1min 2026-07-20 ·