RAZOR: Pruning Replaceable Experts in LLMs
RAZOR, a training-free expert pruning method for mixture-of-experts (MoE) large language models, achieved the highest nine-task macro average among evaluated pruning methods in all eight settings test…
RAZOR, a training-free expert pruning method for mixture-of-experts (MoE) large language models, achieved the highest nine-task macro average among evaluated pruning methods in all eight settings test…
White Circle launched Halo, a post-training framework for open-source models that delivers up to 2.8x the throughput of stock TRL with lower peak memory while keeping models in native HuggingFace form…
Zhipu AI's GLM platform lists GLM-5.3 at 8 yuan per million input tokens, 2 yuan per million cached tokens and 28 yuan per million output tokens with a 1M-token context on its official bigmodel.cn pri…
AkitaOnRails published Part 2 of its LLM Benchmark v4, retesting 39 large language models on a seven-sprint Rails app suite seeded with 14 real CVE-based sabotages, after the author rejected the entir…
A two-stage agentic pipeline combining GLM-4.7-Flash, a 30B tool-calling model, with Poro2, a 70-billion-parameter Finnish-language model, enriches medical terminology via Model Context Protocol tools…
Z.ai (Zhipu) has deprecated its GLM-4.7-Flash model, which will retire on September 10, 2026, six days from now. The model, released on January 19, 2026, supports a 200k context window and 131k max ou…
A Microsoft MVP based in Japan benchmarked seven local large language models on an NVIDIA DGX Spark, running identical Japanese business questions to compare speed, format adherence, and answer qualit…
A new llama.cpp patch called HotPin enables running a 120-billion-parameter mixture-of-experts (MoE) model on just 24GB of RAM, achieving up to 67% memory savings and a 45% speedup over standard swapp…
Zhipu AI's GLM model family spans a 6,000x range in hardware requirements, with the flagship model needing more memory than a typical home has RAM, according to a user who tested GLM-4.7-Flash on a 16…
Researchers at Eternis trained probes on internal representations of large language models to achieve better calibration than chain-of-thought reasoning. The probes detected when models concealed evid…
GLM, the model family from Z.ai (formerly Zhipu AI), is gaining attention for its open weights and strong benchmark numbers. A developer tested GLM-4.7-Flash, a 30B-A3B MoE model, on a single-GPU home…
A team of engineers successfully ran the 754-billion parameter GLM-5.1 large language model on a consumer PC with only 16GB of RAM and a Ryzen 5 5600G CPU, achieving zero crashes or out-of-memory erro…