Is GLM-5.3-Flash Mythos-Level at Cyber? GLM-5.3-Flash matched the cyber-exploitation performance of Claude Mythos Preview on ExploitBench at roughly 6% of the cost, according to a September 2026 blog post by James Mann. Running with a 1 billion token budget per vulnerability, GLM-5.3-Flash achieved Arbitrary Code Execution on 13 of 41 samples versus Mythos' mean of 10, exploiting the v8 package used in Chrome, Edge, and Node.js. The comparison rests on a 500x price gap — GLM-5.3-Flash at $0.25 per million output tokens versus Mythos Preview at $125 per million — which Mann frames as a wake-up call for model evaluators about cheaply accessible cyber capabilities. Back to blog https://generality.org/blog/ September 2026 · By James Mann Is GLM-5.3-Flash Mythos-level at Cyber? We ran GLM-5.3-Flash on ExploitBench with a budget of 1 billion tokens per vulnerability. Flash was able to leverage these tokens into continued performance improvements until eventually it matched the performance of Claude Mythos Preview. As Flash tokens are much cheaper than Mythos tokens, it achieved performance parity at about 6% of the cost. Mythos logs are not published and so an exact curve cannot be found What is ExploitBench? ExploitBench measures an agent’s ability to convert a vulnerability into a series of increasingly severe exploits on well known package v8, used in Chrome, Edge, and Node.js. The scoring can distinguish incremental steps up to severe exploits, such as Arbitrary Code Execution ACE which indicates full control of the target. In our runs, of 41 samples, GLM-5.3-Flash achieves ACE on 13 of them, relative to Mythos’ mean of 10.