cd /news/large-language-models/when-compression-scores-cannot-decid… · home topics large-language-models article
[ARTICLE · art-87143] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=· neutral

When Compression Scores Cannot Decide: Information Boundaries for Group-Robust LLM Pruning

A new arXiv study (2608.02940v1) finds that a dense pruning score with 0.906 split-half reliability predicted a 16.1% gain but selected an endpoint 6.0% and 7.7% worse than two controls, modeling the gap via information interfaces. Across three dense LLMs, a coarse depth allocation cut worst-group perplexity inflation by 12.6–20.9% relative to balanced uniform allocation, and model-specific complete-mask endpoint selection improved by 2.7–8.0%. In OLMoE, router traces predicted singleton direction in 114/192 cases versus 81/192 under the strongest relabeling, with held-out worst-group KL reductions of 13.7% and 7.2%.

read1 min views1 publishedAug 5, 2026

arXiv:2608.02940v1 Announce Type: new Abstract: A reproducible compression statistic can still select the wrong candidate. A dense pruning score with 0.906 split-half reliability predicted a 16.1% gain. Its selected endpoint was 6.0% and 7.7% worse than two controls. We model the gap through information interfaces that delimit which distinctions each statistic supports. For equal-weight groups, a conic law gives the exact pooling price for positive linear fixed-candidate damage, including diagonal and full PSD second moments. Three two-world constructions and an exact observation-fiber radius characterize what pooled moments, group-local moments, and reference-path curvature leave unresolved. A group-resolved diagonal recovers broad damage order (Spearman 0.9239) while fine order remains weak. Relative to balanced uniform allocation, a coarse depth allocation cuts worst-group perplexity inflation by 12.6--20.9% across three dense LLMs. Model-specific complete-mask endpoint selection improves over those references by 2.7--8.0%. In OLMoE, router traces predict singleton direction (114/192 versus 81/192 under the strongest relabeling). Finite-menu decisions on one layer yield held-out worst-group KL reductions of 13.7% and 7.2%. Local measurements construct candidates. Selection is licensed by complete candidate endpoints or a validated uniform guarantee, with uncertainty calibrated to every comparison.

── more in #large-language-models 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/when-compression-sco…] indexed:0 read:1min 2026-08-05 ·