cd /news/artificial-intelligence/z-ai-releases-glm-5-3-beats-fable-5-… · home topics artificial-intelligence article
[ARTICLE · art-96507] src=officechai.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Z.AI Releases GLM 5.3, Beats Fable 5 And GPT 5.6 Sol On CyberBench

Z.ai released GLM-5.3, an open-weights model that scores 84.5 on CyberGym, surpassing Anthropic's Claude Mythos 5 (83.8) and OpenAI's GPT-5.6 Sol (83.6) on the cybersecurity benchmark. The model, built on a 743-billion-parameter base, improved from 4.6 to 28.3 on Terminal-Bench 3.0 and from 46.2 to 66.9 on DeepSWE v1.1, and Z.ai reported finding 2,436 vulnerabilities across 269 real-world projects since GLM-5.2.

read4 min views1 publishedAug 14, 2026
Z.AI Releases GLM 5.3, Beats Fable 5 And GPT 5.6 Sol On CyberBench
Image: Officechai (auto-discovered)

Anthropic and OpenAI had been holding back their models because they felt that they could be misused by China for their cyber capabilities, but now a Chinese model is better than anything dished out by either lab — at least as per one benchmark — for cybersecurity.

Z.ai has released GLM-5.3, the latest update to its open-weights model family, and the headline result is on CyberGym, a benchmark that tests whether a model can find and validate real vulnerabilities starting from white-box source code. GLM-5.3 scores 84.5 on it, edging out both Claude Mythos 5 at 83.8 and GPT-5.6 Sol at 83.6.

Z.ai frames GLM-5.3 as a coding-first release built on the same 743-billion-parameter base as GLM-5.2, with every gain coming from post-training rather than a new architecture. The company says it pushed further on the RL infrastructure it introduced with GLM-5.2 — IndexShare for long-context processing, SAO for long-horizon reinforcement learning, and its open-source training framework slime — and simply threw more environments and compute at the same stack. The coding gains that come out of that are large: GLM-5.3 jumps from 4.6 to 28.3 on Terminal-Bench 3.0, from 46.2 to 66.9 on DeepSWE v1.1, and from 23.8 to 28.5 on Agents’ Last Exam, putting it ahead of GPT-5.6 Sol on that last one. On Z.ai’s own in-house Z.ai Code Bench, the company says GLM-5.3 hits 34.5% completion at its highest effort setting while using notably fewer output tokens than GLM-5.2 needed for a lower score, a token-efficiency gap it says holds up even against Claude Opus 4.8.

GLM 5.3 Benchmarks #

The more striking numbers, though, are on the cyber side, and Z.ai is upfront that they surprised its own researchers. The company says it added vulnerability-discovery data into GLM-5.3’s post-training mix expecting a modest bump, and instead saw the model start reasoning across entire exploitation chains rather than just spotting isolated bugs. On ExploitBench, which grades deeper reasoning about how a real vulnerability could actually be exploited, GLM-5.3 more than doubles GLM-5.2’s score, climbing from 24.4 to 54.4. On ExploitGym, which counts how many exploitation tasks a model can finish inside a fixed time budget, GLM-5.3 completes 105 tasks in two hours and 130 in six, up from 29 and 39 for GLM-5.2. Both of those benchmarks still show GLM-5.3 well behind Mythos 5 and GPT-5.6 Sol, which score in the mid-70s on ExploitBench and complete well over 180 tasks on ExploitGym’s six-hour budget — Z.ai’s own read is that the further up the exploitation chain a test sits, the bigger GLM-5.3’s improvement over its predecessor, and the wider the remaining gap to the closed frontier. CyberGym is the one benchmark where that gap has closed entirely, and GLM-5.3 has actually gone in front.

Z.ai also says it has been quietly running GLM models against real-world codebases with security teams in China since GLM-5.2, and that the exercise has turned up 2,436 vulnerabilities across 269 projects — including kernels, browser engines and network protocols, with the oldest bug reportedly dating back roughly 40 years. The company has built a public ledger tracking which of those findings have been disclosed versus which remain under embargo, with 53 disclosed so far out of the total.

The release lands against a backdrop where Chinese open models finding real security bugs has already become a live story rather than a hypothetical one. Z.ai’s own GLM-5.2 was the model Hugging Face turned to last month when it needed help investigating an autonomous attack traced back to OpenAI’s own systems, after the American models it tried first reportedly struggled to tell an attacker from an incident responder. OpenAI has since pushed out a dedicated cyber-defense model of its own, GPT-5.6-Cyber, through an expanded defender-access program, and Anthropic’s Mythos access remains gated to a small group of vetted partners under its Glasswing initiative even after the export controls on it were lifted at the end of June.

For now, GLM-5.3 is only available through Z.ai’s GLM Coding Plan and its ZCode agent, running on a points-based quota system with off-peak hours priced at half the standard rate. API access and the model weights are set to follow in stages, with Z.ai saying open weights will land roughly two weeks after launch once its own safety evaluation and hardening work is complete.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @z.ai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/z-ai-releases-glm-5-…] indexed:0 read:4min 2026-08-14 ·