Mystery Ox Alpha Model Revealed To Be From Chinese Lab Z.AI Ox Alpha, the anonymous reasoning model that appeared on OpenRouter and OpenCode on August 20, has been revealed as a new iteration of the GLM series from Chinese lab Z.AI, which will release the model's weights tonight. The model, which scored 80% first-pass accuracy on a 10-task DeepSWE benchmark run by developer Ben Davis—ahead of Claude Fable 5's 65% and GPT-5.6 Sol's 52%—was identified through tokenizer fingerprinting that matched GLM on 11 of 11 probes. Z.AI's pricing for GLM 5.3 is $1.40 per million input tokens and $4.40 per million output tokens, with an 81% cache discount, and the model is expected to ship under the MIT license. The model that spent the better part of a week confusing developers, benchmark trackers, and half of AI Twitter has a name attached to it now. As per Bloomberg, Ox Alpha, the free “stealth” model that showed up on OpenRouter and OpenCode on August 20 with no company logo anywhere near it, is a new iteration of the GLM series, the flagship release from Chinese lab Z.AI. Z.AI will release the weights of the model tonight, making the model open for anyone to use. For anyone who missed the last two weeks, Ox Alpha arrived with a listing that told developers everything about what it could do and nothing about who built it. OpenRouter described it as a reasoning model designed for coding, sustained agentic work, and production workloads, and attributed it to “a third-party provider who has chosen to remain anonymous during this preview.” It shipped with a context window just over a million tokens, took text, image, and video as input, and was offered free through both OpenRouter and OpenCode, with OpenCode advertising capacity of 100 trillion tokens a day and Nous Research claiming it could push a quadrillion through its own portal. Within a day of release, developer Ben Davis’s informal 10-task run on the DeepSWE benchmark had Ox Alpha at 80% first-pass accuracy, ahead of Claude Fable 5 at 65% and GPT-5.6 Sol at 52%. That single data point was enough to send the model viral and end DeepSeek’s 56-day streak atop the OpenCode leaderboard. The identity hunt that followed is where the GLM theory actually took shape, and OfficeChai covered both stages of it as they happened. The first piece, published as the model appeared https://officechai.com/ai/stealth-model-ox-alpha-available-for-free-for-a-week-on-openrouter-and-opencode/ , laid out two competing theories: Xiaomi, whose MiMo team had run an identical stealth-launch playbook with MiMo-V2-Pro under the alias Hunter Alpha, and Z.ai, formerly Zhipu AI, based on a tokenizer fingerprinting test that a developer going by “dax” ran across 25 prompts. That test found Ox Alpha’s raw token counts lining up with GLM’s tokenizer far more closely than with any other lab’s. A separate post at the time noted that Ox Alpha stumbled on the same “dirty token” that has historically tripped up Qwen and GLM-family models specifically, which narrowed the field toward a Chinese lab even before anyone had settled on which one. The second piece, a roundup of the theories circulating on X, went further and quantified the tokenizer match: Ox Alpha lined up with GLM on 11 out of 11 tokenizer probes https://officechai.com/ai/which-lab-is-behind-the-viral-ox-alpha-model-these-are-the-theories/ , while no other lab’s model cleared four. DeepSeek, by comparison, burned 98 tokens on a digit probe where GLM used 29. Commentators at the time were already treating Z.ai as the leading candidate, well ahead of a formal announcement. A Chinese lab landing a model that appears competitive with Fable 5 and GPT 5.6 Sol on coding and agentic benchmarks changes the calculus for every company currently pricing frontier access as a premium product. Z.ai’s own pricing for GLM 5.3 sits at $1.40 per million input tokens and $4.40 per million output tokens on its first-party API, with an 81% cache discount bringing cached input down close to a quarter of a cent per thousand tokens, and the model is expected to ship under the same MIT license Z.ai has used since GLM 5. That combination, frontier-adjacent coding performance at a fraction of the cost and released openly, is exactly the pressure point Western labs have spent the last two years trying to avoid. It also explains why Z.ai chose to run the model anonymously on OpenRouter before naming it. A free, unbranded preview lets a lab collect real usage data and impartial benchmark chatter on neutral ground, away from the discount narrative that follows every announcement of an open-weight Chinese model. The cybersecurity numbers are the part worth sitting with longer than the coding scores. A model that can reason across an entire exploitation chain rather than flagging isolated bugs, and that Z.ai says it did not deliberately optimize for, is a capability that spreads the moment the weights go public, expected within roughly two weeks of the announcement. Z.ai has framed this as an emergent property of scaling post-training rather than a deliberate offensive-security push, but that framing does not change who gets access to it once it is open. Every lab building frontier models now has to price in the fact that a capability discovered by accident in one release cycle becomes a baseline expectation across the entire open-weight ecosystem in the next one, and the new GLM model is likely to be cited as the moment that argument stopped being theoretical.