04:00
2026-09-22
machinebrief.com
artificial-intelligence
Improving Parameter Utilization by Sharing Neural Experts Across Layers in Transformers
A new arXiv paper (arXiv:2609.22199v1) proposes CS-MoE, a Transformer architecture that shares neural experts across layers via a centralized global expert pool, achieving lower perplexity than equal-…