Quoting Anthropic Frontier Red Team Anthropic's Frontier Red Team reported that GLM-5.3 achieved full control flow hijacks in 4% of 100 randomly selected tasks from its internal Binary Exploitation benchmark, while Claude Mythos Preview succeeded in 6%. Earlier models, including Claude Opus 4.6 and GLM-5.2, did not succeed on any of the tasks, which Anthropic cited as evidence that a meaningful capability threshold has been crossed. We evaluate several models on 100 tasks from the internal Binary Exploitation benchmark selected at random , and find that GLM-5.3 develops full control flow hijacks in 4% of the trials; Claude Mythos Preview did so in 6%. Although GLM-5.3 performs below Claude Mythos Preview here, a meaningful threshold has clearly been crossed: earlier models, like Claude Opus 4.6 and GLM-5.2, do not succeed in any of them. — Anthropic Frontier Red Team https://www.anthropic.com/research/glm-5-3-and-the-spread-of-advanced-cyber-capabilities , GLM-5.3 and the spread of advanced cyber capabilities Tags: anthropic https://simonwillison.net/tags/anthropic , generative-ai https://simonwillison.net/tags/generative-ai , ai-security-research https://simonwillison.net/tags/ai-security-research , glm https://simonwillison.net/tags/glm , ai https://simonwillison.net/tags/ai , ai-in-china https://simonwillison.net/tags/ai-in-china , llms https://simonwillison.net/tags/llms