Hi everyone,
I am a high school sophomore looking for an arXiv endorser in cs.LG (Machine Learning) or cs.CV (Computer Vision) to help upload my independent preprint.
Brief Overview:
Standard MoE routers discard the self-attention matrix before routing, despite its existing structural information on patch semantics. My paper introduces three scalar signals from this matrix (entropy, cross-head variance, and CLS-similarity) as a routing prior that anneals to zero during training. Through random prior ablations and signal dropout experiments, I found that routing diversity at initialization—rather than semantic content—is what drives early collapse prevention. Semantic attention primarily shapes post-training expert specialization. The work includes 3-seed experiments, four ablation conditions, and honest reporting of a non-significant p > 0.05 result.
Endorsement Criteria:
To endorse this submission, arXiv requires you to have 3+ papers submitted to cs.*
categories within the last 5 years (with at least one paper being older than 3 months).
If you are willing to endorse or have any feedback, please reply here or reach out so I can share the arXiv submission link. Thank you so much for your time and consideration! Best regards,
Hanyu (Olivia) Yu