OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 but only with its own custom test harness
OpenAI claims its GPT-5.6 Sol model scored 38.3 percent on the ARC-AGI-3 benchmark, surpassing Anthropic's Opus 5 at 30.2 percent, but only when using OpenAI's own API with retained reasoning and cont…