Compressing What Matters: Neuron Importance Meets Data-Aware Low Rank Approximation for Language Model Compression
A new study on arXiv proposes combining neuron importance with data-aware low-rank approximation for compressing large language models, achieving performance on par with or better than previous state-…