arXiv:2609.30465v1 Announce Type: cross Abstract: Mixture-of-experts (MoE) models activate few experts per token but store the full expert pool. Expert pruning reduces this storage burden; at a fixed pruning budget, the goal is to preserve the original model's output distribution as closely as possible. Yet an expert's usage or contribution magnitude does not by itself determine the damage caused by its removal. What matters is whether the surviving computation can replace its function. We introduce RAZOR, a training-free expert pruning method that scores functional replaceability using consensus residuals: deviations of expert outputs from the original weighted mixture. An exact single-deletion identity at a fixed layer input accounts for survivor renormalization and router-selected refill, providing local scores aggregated over calibration tokens for budgeted pruning without gradients or recovery training. On GLM-4.7-Flash, Qwen3.6-35B-A3B, DeepSeek-V4-Flash-0731, and Hy3 at 25% and 50% expert removal, RAZOR achieves the highest nine-task macro average among the evaluated pruning methods in all eight settings. On the two backbones with matched REAP benchmark runs, it exceeds REAP by 2.12--5.59 points and wins all 36 paired task comparisons. It also lowers reverse KL relative to REAP in all four matched GLM-4.7-Flash and Qwen3.6-35B-A3B model--budget settings. Analysis of responses generated by Qwen3.6-35B-A3B nevertheless reveals changes in diversity, formatting, and termination, underscoring that task retention and predictive fidelity do not ensure generation stability.
RAZOR: Pruning Replaceable Experts in LLMs
RAZOR, a training-free expert pruning method for mixture-of-experts (MoE) large language models, achieved the highest nine-task macro average among evaluated pruning methods in all eight settings tested on GLM-4.7-Flash, Qwen3.6-35B-A3B, DeepSeek-V4-Flash-0731, and Hy3 at 25% and 50% expert removal, according to the arXiv paper 2609.30465v1. On the two backbones with matched REAP benchmark runs, RAZOR exceeded REAP by 2.12 to 5.59 points and won all 36 paired task comparisons, while also lowering reverse KL relative to REAP in all four matched GLM-4.7-Flash and Qwen3.6-35B-A3B model-budget settings. The paper scores functional replaceability using consensus residuals, deviations of expert outputs from the original weighted mixture, and notes that analysis of Qwen3.6-35B-A3B responses revealed changes in diversity, formatting, and termination, showing task retention and predictive fidelity do not ensure generation stability.
Run your AI side-project on zahid.host
EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.