{"slug": "routing-is-not-enough-diagnosing-intra-adapter-subspace-contention-in-moe-lora", "title": "Routing Is Not Enough: Diagnosing Intra-Adapter Subspace Contention in MoE+LoRA Fine-Tuning", "summary": "A new arXiv paper (2609.03150v1) reports that token-level routing alone does not prevent negative transfer in MoE+LoRA fine-tuning, as adding biomedical data to code and math training increased code perplexity despite near-disjoint expert routing. The authors introduce Jaccard routing overlap and adapter-gradient cosine similarity diagnostics, attributing interference to orthogonal domain gradients competing within the same low-rank adapter subspace, and propose SpawnLoRA, which dynamically adds gated sub-adapters inside MoE experts, reducing negative transfer on Phi-tiny-MoE-instruct and OLMoE-1B-7B compared with standard and rank-adaptive LoRA.", "body_md": "arXiv:2609.03150v1 Announce Type: new\nAbstract: Multi-domain fine-tuning often combines MoE routing with LoRA, assuming that token-level routing separates domain-specific updates. We test this assumption in MoE+LoRA using Python code paired with biomedical text and mathematical reasoning. Although these domains show near-disjoint expert routing, adding biomedical data substantially increases code perplexity, indicating that routing separation alone may not prevent negative transfer. To localize the failure, we introduce Jaccard routing overlap and adapter-gradient cosine similarity, which measure expert sharing and update compatibility, respectively. These diagnostics indicate that interference arises mostly from nearly orthogonal domain gradients competing within the same low-rank adapter subspace. We address this issue with SpawnLoRA, which dynamically adds gated sub-adapters inside MoE experts when adapter-level contention is detected, while keeping the router fixed. We evaluate SpawnLoRA on Phi-tiny-MoE-instruct and OLMoE-1B-7B across multiple mixture settings and find that it effectively reduces negative transfer compared with standard and rank-adaptive LoRA. These results demonstrate that structural separation inside experts provides benefits beyond routing or rank expansion alone.", "url": "https://wpnews.pro/news/routing-is-not-enough-diagnosing-intra-adapter-subspace-contention-in-moe-lora", "canonical_source": "https://arxiv.org/abs/2609.03150", "published_at": "2026-09-04 04:00:00+00:00", "updated_at": "2026-09-04 04:25:28.365734+00:00", "lang": "en", "topics": ["machine-learning", "large-language-models", "ai-research"], "entities": ["arXiv", "Phi-tiny-MoE-instruct", "OLMoE-1B-7B", "SpawnLoRA", "MoE", "LoRA"], "alternates": {"html": "https://wpnews.pro/news/routing-is-not-enough-diagnosing-intra-adapter-subspace-contention-in-moe-lora", "markdown": "https://wpnews.pro/news/routing-is-not-enough-diagnosing-intra-adapter-subspace-contention-in-moe-lora.md", "text": "https://wpnews.pro/news/routing-is-not-enough-diagnosing-intra-adapter-subspace-contention-in-moe-lora.txt", "jsonld": "https://wpnews.pro/news/routing-is-not-enough-diagnosing-intra-adapter-subspace-contention-in-moe-lora.jsonld"}}