{"slug": "resource-efficient-pruning-for-transformer-via-low-rank-importance-estimation", "title": "Resource-Efficient Pruning for Transformer via Low-Rank Importance Estimation", "summary": "Researchers propose REP-LIE, a pruning method that estimates weight importance from LoRA low-rank matrix gradients during finetuning, avoiding full gradient computation and prior finetuning. Tests on LLaMA-7B and Mistral-7B show competitive performance with reduced resource consumption.", "body_md": "arXiv:2608.24973v1 Announce Type: new\nAbstract: With the rapid development of large-scale pre-trained language models based on Transformer architectures, their high computational and memory costs have become a major obstacle to deployment, especially in resource-constrained environments. Traditional pruning methods typically depend on full gradient-based importance estimation, and they necessitate prior finetuning of the model to achieve satisfactory performance. This process often results in intolerable resource consumption. This paper proposes REP-LIE, a new approach to enable resource-efficient pruning during the process of finetuning. REP-LIE leverages the gradients of LoRA low-rank matrices to estimate the importance of weights without requiring full gradient computation. To address the inherent randomness in importance estimation, a stability score is introduced, serving as the basis for iterative pruning of unimportant model parameters. The pruned model is further finetuned through lightweight updates, eliminating the need for full-parameter optimization in the process of finetuning. Extensive experiments on both medium-scale encoder models and large-scale generative models (LLaMA-7B and Mistral-7B) demonstrate that REP-LIE still achieves competitive performance compared to existing approaches.", "url": "https://wpnews.pro/news/resource-efficient-pruning-for-transformer-via-low-rank-importance-estimation", "canonical_source": "https://arxiv.org/abs/2608.24973", "published_at": "2026-08-27 04:00:00+00:00", "updated_at": "2026-08-27 04:19:17.210657+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "ai-research"], "entities": ["REP-LIE", "LoRA", "LLaMA-7B", "Mistral-7B"], "alternates": {"html": "https://wpnews.pro/news/resource-efficient-pruning-for-transformer-via-low-rank-importance-estimation", "markdown": "https://wpnews.pro/news/resource-efficient-pruning-for-transformer-via-low-rank-importance-estimation.md", "text": "https://wpnews.pro/news/resource-efficient-pruning-for-transformer-via-low-rank-importance-estimation.txt", "jsonld": "https://wpnews.pro/news/resource-efficient-pruning-for-transformer-via-low-rank-importance-estimation.jsonld"}}