cd /news/large-language-models/lrcc-generalizing-low-rank-compressi… · home › topics › large-language-models › article
[ARTICLE · art-147320] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=↑ positive

LRCC: Generalizing Low-Rank Compression with Conditional Computation

Low-Rank Conditional Computation (LRCC), a method that trains one lightweight router per Transformer block to select among nested low-rank paths while keeping low-rank factors frozen, improved average downstream accuracy on Llama-2-7B by 7.6 percentage points over static low-rank compression at the same average active-parameter budget, according to the arXiv paper 2610.08858v1. Evaluated on Llama and Qwen models for language modeling and zero-shot tasks, LRCC also improved perplexity and downstream accuracy on Llama-3.2-1B at matched batch-size-1 decoding latency without specialized kernels. The authors additionally analyzed the routers' path choices to assess token-wise path assignment.

by read1 min views1 publishedOct 8, 2026

arXiv:2610.08858v1 Announce Type: new Abstract: Low-rank compression reduces the cost of pretrained language models by replacing linear transformations with low-rank factorizations. However, conventional methods use a fixed rank allocation during inference, assigning the same amount of compute regardless of the input token. We introduce Low-Rank Conditional Computation (LRCC), which adds token-dependent computation to pretrained models by training one lightweight router per Transformer block to select among a small set of nested low-rank paths. During training, the low-rank factors remain frozen, and only the routers are optimized. We evaluate LRCC on Llama and Qwen models for language modeling and zero-shot downstream tasks. Within the same average active-parameter budget, LRCC improves the predictive performance over static low-rank compression, including a 7.6 percentage-point gain in average downstream accuracy on Llama-2-7B over static methods. At matched batch-size-1 decoding latency, LRCC improves both perplexity and downstream accuracy on Llama-3.2-1B and remains competitive on Llama-2-7B, without specialized kernels. Finally, we assess the usefulness of assigning a token-wise path by analyzing the routers' path choices.

── more in #large-language-models 4 stories · sorted by recency
── more on @lrcc 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/lrcc-generalizing-lo…] indexed:0 read:1min 2026-10-08 · —