{"slug": "lrcc-generalizing-low-rank-compression-with-conditional-computation", "title": "LRCC: Generalizing Low-Rank Compression with Conditional Computation", "summary": "Low-Rank Conditional Computation (LRCC), a method that trains one lightweight router per Transformer block to select among nested low-rank paths while keeping low-rank factors frozen, improved average downstream accuracy on Llama-2-7B by 7.6 percentage points over static low-rank compression at the same average active-parameter budget, according to the arXiv paper 2610.08858v1. Evaluated on Llama and Qwen models for language modeling and zero-shot tasks, LRCC also improved perplexity and downstream accuracy on Llama-3.2-1B at matched batch-size-1 decoding latency without specialized kernels. The authors additionally analyzed the routers' path choices to assess token-wise path assignment.", "body_md": "arXiv:2610.08858v1 Announce Type: new \nAbstract: Low-rank compression reduces the cost of pretrained language models by replacing linear transformations with low-rank factorizations. However, conventional methods use a fixed rank allocation during inference, assigning the same amount of compute regardless of the input token. We introduce Low-Rank Conditional Computation (LRCC), which adds token-dependent computation to pretrained models by training one lightweight router per Transformer block to select among a small set of nested low-rank paths. During training, the low-rank factors remain frozen, and only the routers are optimized. We evaluate LRCC on Llama and Qwen models for language modeling and zero-shot downstream tasks. Within the same average active-parameter budget, LRCC improves the predictive performance over static low-rank compression, including a 7.6 percentage-point gain in average downstream accuracy on Llama-2-7B over static methods. At matched batch-size-1 decoding latency, LRCC improves both perplexity and downstream accuracy on Llama-3.2-1B and remains competitive on Llama-2-7B, without specialized kernels. Finally, we assess the usefulness of assigning a token-wise path by analyzing the routers' path choices.", "url": "https://wpnews.pro/news/lrcc-generalizing-low-rank-compression-with-conditional-computation", "canonical_source": "https://arxiv.org/abs/2610.08858", "published_at": "2026-10-08 04:00:00+00:00", "updated_at": "2026-10-08 04:18:16.511778+00:00", "lang": "en", "topics": ["large-language-models", "machine-learning", "ai-research"], "entities": ["LRCC", "Llama-2-7B", "Llama-3.2-1B", "Qwen", "arXiv"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/lrcc-generalizing-low-rank-compression-with-conditional-computation", "markdown": "https://wpnews.pro/news/lrcc-generalizing-low-rank-compression-with-conditional-computation.md", "text": "https://wpnews.pro/news/lrcc-generalizing-low-rank-compression-with-conditional-computation.txt", "jsonld": "https://wpnews.pro/news/lrcc-generalizing-low-rank-compression-with-conditional-computation.jsonld"}}