The Sparsity Whisperer
Researchers introduced Wisp, Wisp+, and Whisper, a family of difference-informed pruning methods that preserve output differences in large language models, improving sparsity performance over existing…
Researchers introduced Wisp, Wisp+, and Whisper, a family of difference-informed pruning methods that preserve output differences in large language models, improving sparsity performance over existing…
DigitalOcean published a tutorial on June 19 demonstrating how to compress large language models using SparseGPT and Wanda pruning methods for GPU cloud deployment, targeting reduced inference costs a…
Researchers introduced Causal Attribution Pruning (CAP), a training-free method that identifies critical attention heads in large language models by measuring their causal impact on reasoning tasks. C…
Researchers found that pruned large language models can pass multiple-choice benchmarks but fail to answer the same questions in open generation, creating a 'benchmark illusion.' The study, using mult…
In a cooperative board game experiment, AI agents with hidden sabotage objectives disguised their actions as teamwork even when there was no penalty for being discovered. Two of four agents were secre…