cd /news/ai-safety/is-glm-5-3-flash-mythos-level-at-cyb… · home topics ai-safety article
[ARTICLE · art-129577] src=generality.org ↗ pub= topic=ai-safety verified=true sentiment=↓ negative

Is GLM-5.3-Flash Mythos-Level at Cyber?

GLM-5.3-Flash matched the cyber-exploitation performance of Claude Mythos Preview on ExploitBench at roughly 6% of the cost, according to a September 2026 blog post by James Mann. Running with a 1 billion token budget per vulnerability, GLM-5.3-Flash achieved Arbitrary Code Execution on 13 of 41 samples versus Mythos' mean of 10, exploiting the v8 package used in Chrome, Edge, and Node.js. The comparison rests on a 500x price gap — GLM-5.3-Flash at $0.25 per million output tokens versus Mythos Preview at $125 per million — which Mann frames as a wake-up call for model evaluators about cheaply accessible cyber capabilities.

read3 min views1 publishedSep 14, 2026
Is GLM-5.3-Flash Mythos-Level at Cyber?
Image: source

Back to blog September 2026 · By James Mann

We ran GLM-5.3-Flash on ExploitBench with a budget of 1 billion tokens per vulnerability. Flash was able to leverage these tokens into continued performance improvements until eventually it matched the performance of Claude Mythos Preview. As Flash tokens are much cheaper than Mythos tokens, it achieved performance parity at about 6% of the cost.

*Mythos logs are not published and so an exact curve cannot be found

What is ExploitBench? #

ExploitBench measures an agent’s ability to convert a vulnerability into a series of increasingly severe exploits on well known package v8, used in Chrome, Edge, and Node.js. The scoring can distinguish incremental steps up to severe exploits, such as Arbitrary Code Execution (ACE) which indicates full control of the target. In our runs, of 41 samples, GLM-5.3-Flash achieves ACE on 13 of them, relative to Mythos’ mean of 10.<sup>1</sup>

Why did you run this experiment? #

Present evaluations generally give all evaluated agents identical budgets in which to work.<sup>2</sup> However, tokens vary enormously in price. When trying to upper bound performance, especially relative to another model, the important comparison is not between tokens but between dollars. GLM-5.3-Flash is priced at $0.25 / M output tokens, Mythos Preview is priced at $125 / M output tokens, a 500x difference.<sup>3</sup> ExploitBench typically would afford them the same number of turns and, predictably, find Flash to be lacking.

This is not the measurement that we actually care about. We’re interested in upper bounding an agents performance, especially relative to other agents, so that we can make inference about the real world impact of systems. It is known that Mythos possesses significant capabilities in cyber, it’s not clear that the world is sufficiently prepared for the democratisation and open-sourcing of those capabilities.

GLM-5.3-Flash may be Mythos level at cyber #

We had to make changes to ExploitBench for it to be possible to run this experiment. Generally, the changes we made were all designed to make it such that runs terminated for one of two reasons, they scored 16/16, or they hit their 1B token caps.<sup>4</sup>

Closing thoughts #

To be clear, we should not infer from this single run that GLM-5.3-Flash is actually at the level of Mythos in every aspect of cyber. There is also a question of contamination, and whether ExploitBench may have been trained against in some way, confounding these results. Nonetheless, assuming that is not the case, this should certainly serve as a wake up call to model evaluators and the broader field. In at least this one case, capabilities significantly above what we believed to be available from our testing were (and are forevermore) accessible to anyone willing to pay the sub $1/hour which a GLM-5.3-Flash agent costs.

  1. Mythos’ 3 runs scored 9, 10, and 11 ACEs, respectively. The reference covers 119 episodes across 40 CVEs. ↩︎
  2. The runs we use to plot other models are from the original implementation of ExploitBench which caps the number of turns, not the number of tokens. This means that while every model shared a maximum of 300 turns, the number of tokens used varies. That said, we think best practice is to use tokens as the more comparable limiting factor. ↩︎
  3. Costs use Flash’s discounted evaluation rates and ExploitBench’s published estimates for Mythos. That means using Mythos Preview’s historic $125 / M output, and GLM-5.3-Flash’s 50% discounted rate of $0.25 / M output. ↩︎
  4. 4/41 samples finished because of errors. These were dealt with by extending the final observed score, that is, pretending these runs never improved, rather than that they crashed. In this way, if anything, we are likely to be underestimating GLM-5.3-Flash’s real capability. ↩︎
── more in #ai-safety 4 stories · sorted by recency
── more on @glm-5.3-flash 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/is-glm-5-3-flash-myt…] indexed:0 read:3min 2026-09-14 ·