GLM-5.3 and the spread of advanced cyber capabilities Anthropic's Frontier Red Team found that Zhipu AI's GLM-5.3, an open-weight model released without meaningful safeguards, can be jailbroken between 64% and 100% of the time with simple techniques, while the same attacks failed against safeguarded Claude models. The assessment, which Anthropic says broadly matches NIST CAISI's Sept. 17 finding that GLM-5.3 is "the most cyber-capable open-weight model released to date" and trails the US frontier by about four months, reports GLM-5.3 built end-to-end exploits in 50 of 410 ExploitBench attempts versus 56 of 410 for Claude Mythos Preview. Anthropic says the lax safeguards significantly expand the cyber capabilities available to malicious actors, though the same capabilities can aid defenders. Subscribe to the Frontier Red Team newsletter Get updates on our latest red-teaming research and findings. Andrew Fasano, Marius Fleischer Cole McFaul, Robert Xiao, Tripp Gallagher Five months ago, we announced https://www.anthropic.com/glasswing Claude Mythos Preview, the first AI model that could autonomously build sophisticated, end-to-end cyber exploits. The rapid rate of improvement in AI suggested to us that this ability would eventually proliferate to many other models, making it much easier for malicious cyber actors to launch highly impactful cyberattacks. In light of these considerations, we chose to release Claude Mythos Preview in a limited way, through Project Glasswing—which enabled trusted cyber defenders to find more than 10,000 vulnerabilities https://www.anthropic.com/research/glasswing-initial-update in critical software, giving them a head start before malicious actors had access to similarly capable models. But those models have now arrived. In this post, we share our analysis of GLM-5.3, the latest AI model developed by Zhipu AI known outside of China as Z.ai . Like Claude Mythos Preview, GLM-5.3 has strong capabilities for autonomously building end-to-end cyber exploits. But GLM-5.3 is unlike other frontier models in that it has been released without meaningful safeguards to limit misuse. We find that attackers can bypass GLM-5.3’s safeguards between 64% and 100% of the time with simple techniques in our simulated tests. In contrast, these attacks did not succeed against safeguarded Claude models in our testing. We assess that GLM-5.3’s lax safeguards significantly increase the cyber capabilities available to malicious actors. At the same time, these capabilities can also benefit defenders working to secure their systems. On Sept. 17, NIST’s Center for AI Standards and Innovation CAISI published its own assessment https://www.nist.gov/news-events/news/2026/09/caisis-assessment-zais-glm-53-cyber-capabilities of GLM-5.3’s cyber capabilities. CAISI found that GLM-5.3 is “the most cyber-capable open-weight model released to date” and that it lags the US frontier by about four months on an aggregate of CAISI’s cyber benchmarks. Our capability findings broadly match CAISI’s. In CAISI’s comparison, US models were tested with cyber safeguards disabled when applicable, and the US frontier includes models released only to vetted users. Attackers can’t readily access those versions of US models, but anyone can download GLM-5.3. This post adds our analysis of how easily GLM-5.3’s safeguards can be bypassed or removed. To understand how GLM-5.3 could enable cyber threat actors to find and exploit real software vulnerabilities, we ran evaluations using automated benchmarks and human-in-the-loop workflows. For both approaches, we ran the tested models in isolated and sandboxed environments so they can only attack offline targets that we have set up for the purposes of these evaluations. We focus primarily on exploit development capability, as this is where Claude Mythos Preview demonstrated a notable jump versus previous Claude models. First, we ran the model on ExploitBench https://arxiv.org/abs/2605.14153 , which measures how well AI models can exploit known vulnerabilities in the V8 engine used by Google Chrome. Here we focus on the models’ ability to develop end-to-end exploits successfully, as this is the most relevant capability for attackers, and where we see significant changes between models. We find that GLM-5.3 develops end-to-end exploits in 50 of 410 attempts. Claude Mythos Preview did so at a similar rate—in 56 of 410 attempts. In our internal Binary Exploitation benchmark,