Expert-Space Exploration in MoE Reinforcement Learning
Recent reinforcement learning advances for Mixture-of-Experts (MoE) large language models have focused on improving optimization stability and training efficiency while treating expert selection as a fixed component, acc…