Z.ai Unveils GLM-5.3 for Coding and Cybersecurity Z.ai unveiled GLM-5.3 on August 14, reporting improved performance on long-horizon coding and software-vulnerability analysis. The model scored 84.5% on Z.ai's CyberGym benchmark, above the company's reported 83.8% for Anthropic's Mythos 5, but trailed on exploit-development tests with 54.4% versus 78.0%. Z.ai said public weights would follow a two-week security review, with sensitive cyber functions restricted to verified users. Z.ai Unveils GLM-5.3 for Coding and Cybersecurity Z.ai unveiled GLM-5.3 on August 14, reporting improved performance on long-horizon coding and software-vulnerability analysis. Reuters reports that the model scored 84.5% on Z.ai's CyberGym benchmark, above the company's reported 83.8% result for Anthropic's Mythos 5, while trailing on exploit-development tests. Z.ai said public weights would follow a two-week security review, with sensitive cyber functions restricted to verified users. Z.ai unveiled GLM-5.3 on August 14, presenting the model as an upgrade for complex coding, long-running agent tasks, and cybersecurity analysis. Reuters reports that Z.ai intended to make the model publicly available after about two weeks of security assessments and safeguard hardening, while initially providing access to a limited set of launch partners. The release is not a new base-model pretraining run. MarkTechPost reports that GLM-5.3 uses the same 743 billion-parameter base model as GLM-5.2, with the reported improvements coming from expanded post-training. Interconnects similarly reports that Z.ai described the work as substantially extended post-training rather than a change to the underlying base model. Reported coding gains According to MarkTechPost's account of Z.ai's results, GLM-5.3 improved from 4.6 to 28.3 on Terminal-Bench 3.0, from 46.2 to 66.9 on DeepSWE v1.1, and from 23.8 to 28.5 on Agents' Last Exam CLI, relative to GLM-5.2. These are vendor-reported benchmark figures, and comparisons can depend materially on the task harness, context limits, sampling configuration, tool setup, and token budget. MarkTechPost reports that GLM-5.3 was available through Z.ai's API, GLM Coding Plan, and ZCode, although the model weights were not yet public. For engineering teams, the distinction between hosted access and weight availability matters: local deployment, controlled fine-tuning, and independent security evaluation require a different operational path than API adoption. Long-horizon coding evaluations are useful because they test more than code completion. They can require an agent to inspect repositories, execute commands, diagnose failures, preserve state across steps, and validate a final change. Across the industry, stronger results in such environments do not by themselves establish reliability on production repositories, where permission boundaries, CI systems, proprietary dependencies, and rollback processes are central constraints. Cybersecurity capability and access controls Reuters reports that Z.ai scored 84.5% on CyberGym, a benchmark for reviewing code, identifying vulnerabilities, and confirming that flaws are genuine. Z.ai's reported comparison placed Anthropic's Mythos 5 at 83.8%, but Reuters noted that the results had not been independently verified. The South China Morning Post reported the same Z.ai figure and additionally cited a company comparison of 83.6% for OpenAI's GPT-5.6 Sol. The CyberGym result is only one part of the security picture. Reuters reports that GLM-5.3 scored 54.4% on ExploitBench, compared with Z.ai's reported 78.0% for Mythos 5. In timed attack-development tasks cited by Reuters, Z.ai reported 105 completed tasks in two hours and 130 in six hours, versus 181 and 247 for Mythos 5. That split is significant for security practitioners. Vulnerability discovery and validation can support secure-code review, triage, and defensive testing, while generating working attacks can introduce more direct misuse risk. Reuters reports that Z.ai would place its most sensitive cybersecurity features behind a "trusted access" program for verified users. Comparable frontier-model releases increasingly separate broad model availability from access to high-risk tooling, testing environments, or capability configurations. What remains to validate The published results indicate a sizable vendor-reported post-training improvement, particularly on agentic coding benchmarks. Independent replication, detailed evaluation harnesses, and security testing by external teams remain important for assessing how the results transfer to real repositories and defensive workflows. Reuters reported that Anthropic restricts Mythos, a version of Claude Fable 5 with cybersecurity safeguards removed, to vetted organizations. That comparison also underscores why benchmark leadership in vulnerability identification should be read alongside access controls, auditability, and red-team evidence rather than as a stand-alone deployment recommendation. Key Points - 1Z.ai reports substantial GLM-5.3 gains from post-training alone, reinforcing the importance of post-training data, environments, and evaluation design. - 2The model's CyberGym lead and ExploitBench deficit separate vulnerability discovery from end-to-end exploitation capability, a crucial distinction for security teams. - 3Restricted access to sensitive cyber functions reflects a broader industry pattern of pairing advanced security models with identity and safety controls. Scoring Rationale GLM-5.3 is a notable open-weight-model development with reported gains in agentic coding and vulnerability analysis, two high-value workflows for ML and software teams. Its security capability claims and phased-access approach make the release relevant beyond routine benchmark competition, although the central results remain vendor-reported. Sources Public references used for this report. Practice interview problems based on real data 1,625 SQL & Python problems across 15 industry datasets — the exact type of data you work with. Try 250 free problems /problems