We evaluate several models on 100 tasks from the [internal Binary Exploitation benchmark] (selected at random), and find that GLM-5.3 develops full control flow hijacks in 4% of the trials; Claude Mythos Preview did so in 6%. Although GLM-5.3 performs below Claude Mythos Preview here, a meaningful threshold has clearly been crossed: earlier models, like Claude Opus 4.6 and GLM-5.2, do not succeed in any of them.
— [Anthropic Frontier Red Team](https://www.anthropic.com/research/glm-5-3-and-the-spread-of-advanced-cyber-capabilities), GLM-5.3 and the spread of advanced cyber capabilities
Tags: [anthropic](https://simonwillison.net/tags/anthropic), [generative-ai](https://simonwillison.net/tags/generative-ai), [ai-security-research](https://simonwillison.net/tags/ai-security-research), [glm](https://simonwillison.net/tags/glm), [ai](https://simonwillison.net/tags/ai), [ai-in-china](https://simonwillison.net/tags/ai-in-china), [llms](https://simonwillison.net/tags/llms)