{"slug": "pushing-gpt-5-6-sol-from-13-3-to-100-on-arc-agi-3-public", "title": "Pushing GPT-5.6-Sol from 13.3% to 100% on ARC-AGI-3 Public", "summary": "INT21's SwarmOS infrastructure orchestration platform achieved a 100% RHAE score on the ARC-AGI-3 Public benchmark using only GPT-5.6-Sol, up from the model's official 13.3%, a 7.5x improvement, and also boosted GPT-5.6-Luna's score from 0% to 56%. The results, which cover the 25-environment public set, suggest that system architecture can significantly enhance model capabilities, with SwarmOS solving 25 environments and 183 levels in 6,731 actions, compared to NVIDIA AVO's 6,624 and VISTA's 7,542 actions using Claude Opus 5. INT21 believes this is an early sign of AGI and that the unit of intelligence is shifting from the model to the system.", "body_md": "**For long-horizon agents, model capability alone does not determine system\ncapability. Infrastructure orchestration multiplies what models can do.**\n\nINT21’s SwarmOS tested this hypothesis on the ARC-AGI-3 benchmark and found that infrastructure can dramatically unlock a model’s capabilities.\n\n## Key Takeaways\n\n- SwarmOS achieved a 100% RHAE (Relative Human Action Efficiency) score on ARC-AGI-3 Public using only GPT-5.6-Sol. This is the first time GPT-5.6 has matched Anthropic models on an ARC-AGI-3-style benchmark with action-efficiency parity. The result shows that infrastructure architecture is the differentiator.\n- We also tested GPT-5.6-Luna, the least expensive frontier model in this comparison, with SwarmOS and improved its RHAE score from 0% to 56%. Model intelligence and agent intelligence are not the same thing.\n- SwarmOS is a large-scale, cloud-native evolution infrastructure for generating\nsystem software, including\n[faster video, audio, and music inference compared with SGLang and vLLM](/insights/addressing-the-inference-bottleneck/). Without any ARC-specific specialization, it achieved a full score on ARC-AGI-3 Public. We believe the unit of intelligence is shifting from the model to the system, and that this is an early sign of AGI. - For enterprises budgeting for expensive frontier models, this result suggests that cheaper models paired with superior orchestration can deliver strong long-horizon performance.\n\n## Results\n\n| Model | Official | INT21 SwarmOS | API Price (Input/Output) |\n|---|---|---|---|\n| GPT-5.6-Sol | 13.3% | 100% | $4.0/$20.0 |\n| GPT-5.6-Luna | 0.00% | 56% | $0.2/$1.2 |\n| Claude Opus 5 | 30.16% | No Support | $5.0/$25.0 |\n\nSwarmOS reached a full RHAE score on ARC-AGI-3’s public set with Sol, a 7.5x improvement. With Luna, RHAE improved from 0% to 56% despite Luna being a substantially smaller and less expensive model. This suggests that long-horizon intelligence is not solely a property of the foundation model; system architecture can contribute a meaningful share of the capability.\n\nSwarmOS with GPT-5.6-Sol solved 25 environments and 183 levels, taking 6,731\nenvironment actions to achieve a 100% RHAE score. For reference,\n[NVIDIA’s AVO research project](https://arxiv.org/html/2603.24517v1), co-authored\nby INT21 founder Bing Xu earlier this year, reported 6,624 environment actions.\nVISTA reported 7,542 environment actions to reach the same score on the same\npublic set. Both NVIDIA AVO and VISTA are powered by Anthropic Claude Opus 5.\nThese results cover the 25-environment ARC-AGI-3 public set; they are not results\nfrom the semi-private or fully private competition sets.\n\nWe believe SwarmOS achieving a full score on ARC-AGI-3 is an early sign of AGI (artificial general intelligence). The self-improving agent swarms override learned patterns when they reach dead ends, preventing them from drifting toward globally suboptimal systems. This means orchestration systems can work across very different domains, from pixel games to complex infrastructure generation. The ability to challenge assumptions and change direction when evidence demands it mirrors an aspect of human general intelligence.\n\n## The NVIDIA AVO Context\n\nNVIDIA AVO demonstrated that an orchestration harness can unlock frontier\nperformance. NVIDIA connected Claude Opus 5 to the AVO Harness and reached 100%\non ARC-AGI-3, a 3.3x improvement. [SwarmOS](/#swarmos), our cloud-native\nplatform for running self-improving agents, builds on this finding and tests a\nmore challenging setting: using only GPT-5.6-Sol from a 13.3% baseline, we\nachieved 100%, a 7.5x improvement, with action-efficiency parity against Claude\nOpus 5-based solutions.\n\n## What We Tested\n\n[ARC-AGI-3](https://arcprize.org/tasks) is a benchmark that measures how well AI\nagents learn and reason through unfamiliar, game-like environments. Agents must\ninfer how the environments work without explicit instructions. They succeed by\npreserving useful knowledge, learning from failures, and progressing across\ndozens of challenges without human intervention.\n\nFor this evaluation, we connected the same SwarmOS that powers our system-programming work to the ARC-AGI-3 task interface, with web search turned off and no ARC-specific specialization. We tested two models: GPT-5.6-Sol and GPT-5.6-Luna. SwarmOS coordinates self-improving agent swarms to explore solutions in parallel, share what they learn, and continuously improve results. The architecture was the same across both models; only the underlying LLM changed.\n\n## Why This Matters\n\nNVIDIA’s AVO research showed that an orchestration harness can shift a frontier model, Opus 5, from 30% to 100% on ARC-AGI-3. SwarmOS independently validates this finding and demonstrates a 0% to 56% leap with GPT-5.6-Luna.\n\nThis changes the competitive conversation because model performance is no longer the only constraint. AI users can stop asking only, “Which frontier model should we use?” and start asking, “Which orchestration layer maximizes our ROI?” The discussion now focuses on infrastructure as the new moat.\n\n## What’s Next\n\nThe industry narrative to date has centered on model capability and which company builds the best frontier model. SwarmOS and AVO suggest a different question for the next phase: “Which organization orchestrates its models most effectively?”\n\nWhile competition between frontier models has been fierce, we believe an equally exciting battle in infrastructure is emerging. Enterprises that treat orchestration as a core engineering competency, rather than a bolt-on integration layer, will have a clear competitive advantage in self-improving infrastructure. SwarmOS is purpose-built for this shift: a platform where self-improving infrastructure becomes your competitive moat.", "url": "https://wpnews.pro/news/pushing-gpt-5-6-sol-from-13-3-to-100-on-arc-agi-3-public", "canonical_source": "https://int21.ai/insights/pushing-gpt-5-6-sol-to-100-on-arc-agi-3-public/", "published_at": "2026-08-27 00:00:00+00:00", "updated_at": "2026-08-27 14:51:10.039607+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-agents", "ai-research", "ai-infrastructure"], "entities": ["INT21", "SwarmOS", "GPT-5.6-Sol", "GPT-5.6-Luna", "ARC-AGI-3", "Claude Opus 5", "NVIDIA AVO", "VISTA"], "alternates": {"html": "https://wpnews.pro/news/pushing-gpt-5-6-sol-from-13-3-to-100-on-arc-agi-3-public", "markdown": "https://wpnews.pro/news/pushing-gpt-5-6-sol-from-13-3-to-100-on-arc-agi-3-public.md", "text": "https://wpnews.pro/news/pushing-gpt-5-6-sol-from-13-3-to-100-on-arc-agi-3-public.txt", "jsonld": "https://wpnews.pro/news/pushing-gpt-5-6-sol-from-13-3-to-100-on-arc-agi-3-public.jsonld"}}