ZCode: Z.ai's coding agent harness, explained
Z.ai's ZCode, an Apache-2.0 coding agent workspace with desktop, browser and terminal interfaces, had 7,541 GitHub stars as of 8 Oct 2026 and one release, v3.14.3, published 24 Sep 2026, according to …
Z.ai's ZCode, an Apache-2.0 coding agent workspace with desktop, browser and terminal interfaces, had 7,541 GitHub stars as of 8 Oct 2026 and one release, v3.14.3, published 24 Sep 2026, according to …
Cantina Security, with Yeta Labs, released apex-flash-1, an open-weights vulnerability-research model built as a reinforcement learning fine-tune of Z.ai's GLM-5.3-Flash, solving 40 of 60 held-out bug…
Cantina released apex-flash-1, an open-weights cybersecurity model fine-tuned from GLM-5.3-Flash with 321 billion total parameters, in an October 4th post by co-founder and CEO Harikrishnan "Hari" Mul…
Privatemode published a method on 24 September 2026 that turns an ordinary LLM into a typed decision-maker by forcing one output token and reading its logprobs, scoring 10 wins, 8 ties and 10 losses a…
Inception released Mercury Voice, a diffusion LLM tuned for voice agents, to general availability for enterprise customers, claiming a median (p50) time to first answer token of under 320 milliseconds…
Inception made Mercury Voice generally available to enterprise customers on September 29th, claiming a company-reported 320-millisecond median time to first answer token on production customer-service…
INT21 generated 20 inference engines across seven model categories in two weeks using Rust with C++ and CUDA components, and its MiMo engine reached 1,308 tokens/s versus 540 for tuned SGLang and 1,01…
A developer reported that constraining GLM-5.3-Flash's prompt to end with "Result: " forces a single-token decision output that matches the specialized Jev model on accuracy and latency when run throu…
Privatemode researchers Johannes Hötter and Marko Rosenmüller demonstrated that the off-the-shelf LLM GLM-5.3-Flash can make typed decisions with a probability for every option in a single forward pas…
A September 2026 comparison of open-source coding models found GLM-5.3-Flash from Z.ai best for agentic coding, DeepSeek V4 Flash cheapest per token, and MiniCPM5-2B best for on-device use. GLM-5.3-Fl…
Z.ai released GLM-5.3-Flash on August 26, 2026, a 320-billion-parameter open-weight Mixture-of-Experts model that activates only 18 billion parameters per token and supports a one-million-token contex…
Digital Applied published a six-signal taxonomy on September 21, 2026 for determining whether an AI model launch is imminent, requiring each signal to be verifiable via a public URL and a named field.…
Z.ai says an agent running on its GLM-5.3 model built the inference service that now serves GLM-5.3-Flash, reaching 3.22x baseline throughput in thirteen days on a cluster of more than 100,000 Chinese…
Z.ai released GLM-5.3-Flash and GLM-5.3-FlashX, the first native multimodal models in the GLM-5 series, with GLM-5.3-FlashX delivering inference speeds of 200 tokens/s. The models accept video, image,…
Z.ai released GLM-5.3-FlashX, a native multimodal model that delivers inference speeds of up to 200 tokens/s, faster than GLM-5.3-Flash. The model uses a hybrid sparse and linear attention architectur…
Zhipu AI's GLM platform lists GLM-5.3 at 8 yuan per million input tokens, 2 yuan per million cached tokens and 28 yuan per million output tokens with a 1M-token context on its official bigmodel.cn pri…
Z.AI has named GLM-6.0 and placed "Full Self-Training" at the center of its roadmap, describing a loop that spans pre-training, mid-training, and post-training with self-generated experience, evaluati…
GLM-5.3-Flash matched the cyber-exploitation performance of Claude Mythos Preview on ExploitBench at roughly 6% of the cost, according to a September 2026 blog post by James Mann. Running with a 1 bil…
Void Linux maintainer Andrea Brancaleoni orphaned more than 100 packages he maintained after a dispute over the project's AI usage policy, which requires all AI tool usage to be disclosed in contribut…
A developer analysis of Artificial Analysis benchmarking data finds that GLM-5.3-Flash scores 42 at $0.25 per task while Kimi K3 scores 44 at $2.00 per task, an eight-fold cost gap for a two-point cap…