{"slug": "from-experts-to-sub-experts-fine-grained-parameter-efficient-fine-tuning-for-moe", "title": "From Experts to Sub-experts: Fine-grained Parameter-Efficient Fine-Tuning for MoE LLMs", "summary": "Researchers proposed NSFT (Neural Sub-expert Fine-Tuning), a fine-grained parameter-efficient fine-tuning framework that refines Mixture-of-Experts (MoE) large language model adaptation from experts to sub-experts by decomposing each expert along the intermediate dimension into structured channel groups and selecting task-relevant sub-experts via routing importance combined with intra-expert activation saliency. Experiments on OLMoE and Ling-mini-2.0 across domain-specific tasks and general benchmarks showed NSFT consistently outperformed representative PEFT and expert-level sparse tuning baselines while using substantially fewer trainable parameters and preserving competitive general capability. The authors conclude that sub-expert-level adaptation is a more precise and efficient PEFT paradigm for MoE LLMs.", "body_md": "arXiv:2609.25655v1 Announce Type: new \nAbstract: As large language models (LLMs) scale rapidly, dense full-parameter adaptation becomes increasingly expensive, motivating sparse and modular architectures such as Mixture-of-Experts (MoE) models. This shift raises a key question for parameter-efficient fine-tuning (PEFT): at what granularity should parameters be selected and updated? Existing PEFT methods such as LoRA operate on predefined weight matrices, while expert-level sparse tuning methods update entire selected experts. However, we observe that activated experts are internally sparse, with only a small fraction of intermediate channels strongly responding to downstream tasks, indicating that expert-level adaptation is still too coarse. We propose NSFT (Neural Sub-expert Fine-Tuning), a fine-grained PEFT framework that refines MoE adaptation from experts to sub-experts. NSFT decomposes each expert along the intermediate dimension into structured channel groups and selects task-relevant sub-experts by combining routing importance with intra-expert activation saliency. To optimize sparse partial updates, NSFT further introduces learning-rate scaling and dynamic gradient scaling to compensate for the reduced effective update magnitude. Experiments on OLMoE and Ling-mini-2.0 across challenging domain-specific tasks and general benchmarks show that NSFT consistently outperforms representative PEFT and expert-level sparse tuning baselines, while using substantially fewer trainable parameters and preserving competitive general capability. These results suggest that sub-expert-level adaptation is a more precise and efficient PEFT paradigm for MoE LLMs.", "url": "https://wpnews.pro/news/from-experts-to-sub-experts-fine-grained-parameter-efficient-fine-tuning-for-moe", "canonical_source": "https://www.machinebrief.com/news/from-experts-to-sub-experts-fine-grained-parameter-efficient-v3sz", "published_at": "2026-09-23 04:00:00+00:00", "updated_at": "2026-09-23 04:54:55.283186+00:00", "lang": "en", "topics": ["machine-learning", "large-language-models", "ai-research", "natural-language-processing"], "entities": ["NSFT", "OLMoE", "Ling-mini-2.0", "LoRA", "Mixture-of-Experts"], "alternates": {"html": "https://wpnews.pro/news/from-experts-to-sub-experts-fine-grained-parameter-efficient-fine-tuning-for-moe", "markdown": "https://wpnews.pro/news/from-experts-to-sub-experts-fine-grained-parameter-efficient-fine-tuning-for-moe.md", "text": "https://wpnews.pro/news/from-experts-to-sub-experts-fine-grained-parameter-efficient-fine-tuning-for-moe.txt", "jsonld": "https://wpnews.pro/news/from-experts-to-sub-experts-fine-grained-parameter-efficient-fine-tuning-for-moe.jsonld"}}