# GLM 5.3 Scores 60 On AA Intelligence Index, Three Chinese Labs Now Right Behind US Frontier

> Source: <https://officechai.com/ai/glm-5-3-scores-60-on-aa-intelligence-index-three-chinese-labs-now-right-behind-us-frontier/>
> Published: 2026-08-19 07:55:07+00:00

Chinese labs are building pressure in numbers behind the US frontier.

Z.AI’s GLM-5.3 has landed at 60 on the Artificial Analysis Intelligence Index, tying [Kimi K3](https://officechai.com/ai/kimi-k3-2-8-trillion-parameters-pricing-context-window/) for the top spot among open-weights models and sitting just three points behind Claude Opus 5’s 63. Seven points up from GLM-5.2, the jump is the kind of single-generation gain that used to be rare and but is now becoming routine for some AI labs. Once the weights are out, which Z.AI says should happen within the week, GLM-5.3 will be tied as the leading open-weights model in the world alongside Kimi K3, a milestone that would have sounded implausible even a year ago when Chinese open models were still playing catch-up rather than trading blows with the frontier.

What stands out about this release is not the headline number so much as where the gains actually came from. GLM-5.3’s biggest leap is in agentic capability, the kind of work that involves stringing together tool calls, navigating real software environments, and completing multi-step tasks without constant supervision. On GDPval-AA v2, Artificial Analysis’s real-world agentic knowledge work evaluation, GLM-5.3’s Elo rating climbs from 1524 to 1770. That’s a 246-point jump in one release, and it puts the model second on that particular test behind only Claude Opus 5’s 1855, ahead of every other model on the board including GPT-5.6 Sol. It also clears the previous open-weights leader on GDPval-AA v2, Kimi K3 at 1668, by more than 100 points. For a model that kept the same 753 billion total and 40 billion active parameter architecture as its predecessor, that’s an unusually large capability gain to extract purely from post-training.

The catch is that GLM-5.3 got noticeably less efficient in getting there. Across the Intelligence Index’s nine evaluations, the model uses around 18,700 output tokens per task on average, up roughly 20% from GLM-5.2’s 15,700 and 27% more than what Kimi K3 needs. More reasoning tokens generally means more cost, and that shows up directly in Artificial Analysis’s pricing data. GLM-5.3 costs $0.68 per Intelligence Index task, 1.5 times what GLM-5.2 charged at $0.44. Even so, that figure still undercuts the competition sitting in the same intelligence bracket. GLM-5.3 comes in 19% cheaper than Kimi K3’s $0.84 per task and 45% cheaper than GPT-5.6 Sol’s $1.23, which keeps the model comfortably on the Pareto frontier despite the token bloat.

The other meaningful shift is in factual reliability. GLM-5.3 scores 14 on AA-Omniscience, Artificial Analysis’s measure of real-world knowledge accuracy, up from just 4 for GLM-5.2. That places it second among open-weights models on the metric, still behind Kimi K3’s 20 but a real step forward from where GLM-5.2 stood. The breakdown behind that number matters as much as the number itself. Accuracy rate rose from 24% to 34% and attempt rate climbed from 46% to 55%, which suggests the model is actually answering more questions correctly rather than simply getting better at declining to answer when unsure. The one blemish is that GLM-5.3’s hallucination rate ticked up slightly, from 26% to 30%, a tradeoff that comes with a model willing to attempt more.

On the technical side, GLM-5.3 keeps its predecessor’s footprint of 753 billion total parameters with 40 billion active in a mixture-of-experts setup, paired with a 1 million token context window. Pricing on Z.AI’s first-party API sits at $1.40 per million input tokens and $4.40 per million output tokens, with an 81% cache hit discount that brings cached input down to $0.26 per million tokens. The model ships under the MIT license, consistent with Z.AI’s approach across the [GLM family](https://officechai.com/ai/a-chinese-open-source-model-is-ahead-of-all-google-models-on-the-artificial-analysis-intelligence-index-for-the-first-time/) since GLM-5, though the weights themselves are not out yet.

Zoom out and the shape of the leaderboard tells its own story. Behind Claude Opus 5, Claude Fable 5, and GPT-5.6 Sol at the very top, the next tier is no longer dominated by a single Chinese challenger picking off proprietary models one at a time. It’s three of them clustered together. GLM-5.3 and Kimi K3 sit tied at 60, and [DeepSeek’s V4 Flash 0731](https://officechai.com/ai/deepseek-v4-flash-0731-scores-50-on-artificial-analysis-intelligence-index-creates-big-spike-on-pareto-frontier/) holds its own further down the chart at 52, ahead of Gemini 3.6 Flash at the same score. Z.AI, Moonshot, and DeepSeek are no longer isolated data points climbing the chart in sequence. They’re arriving together, on overlapping release schedules, each pushing a different part of the stack, agentic reasoning here, cost efficiency there, knowledge accuracy somewhere else. For the American labs still holding the top three slots, the gap to any one Chinese model might look manageable in isolation. The gap to three of them moving in formation is a different problem entirely.
