We Raced Seven AI Models to RCE
In a controlled race, OpenAI's GPT 5.6 Sol achieved verified remote code execution in 9 of 11 vulnerable-target runs, the most among seven AI models tested, while Zhipu AI's GLM 5.3 verified 8 runs. The test, which used …