Zhipu launches GLM-5.3-Flash, its first natively multimodal model built for Chinese chips Z.ai, formerly Zhipu AI, released GLM-5.3-Flash on August 26, its first natively multimodal model in the GLM-5 series, running entirely on domestically produced Chinese AI accelerators. The 320-billion-parameter model activates only 18 billion per token, has a 1 million token context window, and scored 63.4 on the DeepSWE v1.1 coding benchmark, up from 46.2 for GLM-5.2. API pricing is $0.15 per million input tokens and $0.50 per million output tokens, with cached inputs at $0.03, and the model is available under the MIT license on Hugging Face and ModelScope. Photo: ed br / Pexels Zhipu launches GLM-5.3-Flash, its first natively multimodal model built for Chinese chips The 320-billion-parameter model runs entirely on domestic AI accelerators and costs a fraction of comparable proprietary alternatives. China’s Z.ai, the company formerly known as Zhipu AI, released GLM-5.3-Flash on August 26, making it the first natively multimodal model in the GLM-5 series. The launch is notable for reasons beyond the model itself: it runs entirely on domestically produced AI accelerators, a pointed demonstration that Chinese hardware can handle serious large-scale workloads. The timing matters. Export controls from the US have restricted access to NVIDIA’s most advanced chips for Chinese buyers, pushing companies like Z.ai to build around what they have. What the model actually does GLM-5.3-Flash uses a hybrid architecture combining sparse and linear attention, which is a design choice that cuts compute requirements significantly. The model carries 320 billion total parameters but activates only 18 billion per token, meaning it draws on a massive knowledge base without paying the full computational cost on every inference. The context window sits at 1 million tokens, placing it among the longest available in any open model. Native support for image and video inputs makes it genuinely multimodal from the architecture level up, rather than as a bolted-on capability. On the DeepSWE v1.1 coding benchmark, GLM-5.3-Flash scored 63.4, compared to 46.2 for its predecessor GLM-5.2. The model also scored 57 on the Artificial Analysis Intelligence Index. Pricing is aggressive. API access costs $0.15 per million input tokens and $0.50 per million output tokens, with cached inputs available at $0.03. Z.ai describes the cost as roughly one-tenth of some competing options. Chinese chips doing real work Z.ai deployed GLM-5.3-Flash across clusters of domestically manufactured AI accelerators, then built custom serving stacks and applied quantization techniques tailored to that hardware. The result was a threefold improvement in end-to-end performance compared to a more generic deployment approach. The model ships with open weights under the MIT license, available in both FP8 and BF16 formats on Hugging Face and ModelScope. MIT is about as permissive as licenses get: developers and companies can use, modify, and distribute the weights commercially without restriction. Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy https://cryptobriefing.com/editorial-policy/ .