OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 but only with its own custom test harness OpenAI claims its GPT-5.6 Sol model scored 38.3 percent on the ARC-AGI-3 benchmark, surpassing Anthropic's Opus 5 at 30.2 percent, but only when using OpenAI's own API with retained reasoning and context compaction. In the official test environment, GPT-5.6 Sol managed just 7.8 percent, while Opus 5 achieved its score without such aids. OpenAI counters Anthropic's ARC-AGI-3 record: GPT-5.6 Sol scores 38.3 percent, but only through its own API with retained reasoning and context compaction. In the official test environment, the model managed just 7.8 percent. Opus 5 hit its 30.2 percent without such aids. The article OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 but only with its own custom test harness https://the-decoder.com/openai-claims-gpt-5-6-sol-beats-opus-5-on-arc-agi-3-but-only-with-its-own-custom-test-harness/ appeared first on The Decoder https://the-decoder.com .