{"slug": "the-sparsity-whisperer", "title": "The Sparsity Whisperer", "summary": "Researchers introduced Wisp, Wisp+, and Whisper, a family of difference-informed pruning methods that preserve output differences in large language models, improving sparsity performance over existing baselines. Across Llama 2 and 3.1 models from 7B to 405B parameters, the second-order Whisper method consistently outperformed reconstruction-based baselines, while update-free variants improved over activation-aware methods like Wanda and SparseGPT, extending to structured sparsity and other model families.", "body_md": "arXiv:2608.06630v1 Announce Type: new\nAbstract: Pruning reduces the inference cost of large language models, but existing criteria primarily preserve large activations or reconstruct layer outputs. We argue that this overlooks a key computation performed by particularly sparsity-sensitive neurons in the MLP up and gate projections: separating similar inputs into dissimilar outputs. This suggests that effective pruning should preserve not only activations, but also the differences between outputs more broadly. We introduce a family of difference-informed pruning methods built upon this principle. Wisp is a first-order, update-free method that scores weights using input-difference norms, and Wisp+ refines this score neuronwise using the input pairs each neuron separates most strongly. Finally, Whisper is a second-order method that uses a lightly regularized difference Hessian as its reconstruction objective. Across Llama 2 and 3.1 models from 7B to 405B parameters, our second-order variant consistently improves over strong reconstruction-based baselines, while our update-free variants improve over activation-aware baselines, especially in constrained settings. The improvements over Wanda and SparseGPT extend to structured sparsity, downstream evaluations, and other model families. Augmenting stronger techniques such as RIA and ALPS with our difference-informed criteria yields further improvements, shifting the overall accuracy-runtime frontier outward at negligible additional cost. These results suggest that preserving output differences is a broadly useful and composable signal for post-training LLM sparsification.", "url": "https://wpnews.pro/news/the-sparsity-whisperer", "canonical_source": "https://www.machinebrief.com/news/the-sparsity-whisperer-eoic", "published_at": "2026-08-10 04:00:00+00:00", "updated_at": "2026-08-10 05:11:53.190757+00:00", "lang": "en", "topics": ["large-language-models", "machine-learning", "artificial-intelligence"], "entities": ["Wisp", "Wisp+", "Whisper", "Llama 2", "Llama 3.1", "Wanda", "SparseGPT", "RIA"], "alternates": {"html": "https://wpnews.pro/news/the-sparsity-whisperer", "markdown": "https://wpnews.pro/news/the-sparsity-whisperer.md", "text": "https://wpnews.pro/news/the-sparsity-whisperer.txt", "jsonld": "https://wpnews.pro/news/the-sparsity-whisperer.jsonld"}}