04:00
2026-07-23
arxiv.org
artificial-intelligence
AdaRoPE: Not All Attention Heads Should Rotate and Scale Equally
A new study from arXiv introduces AdaRoPE, a method that assigns learnable rotation frequencies and attention scaling factors to each attention head in Transformers, outperforming standard Rotary Posi…