cd /news/artificial-intelligence/adarope-not-all-attention-heads-shou… · home topics artificial-intelligence article
[ARTICLE · art-69593] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

AdaRoPE: Not All Attention Heads Should Rotate and Scale Equally

A new study from arXiv introduces AdaRoPE, a method that assigns learnable rotation frequencies and attention scaling factors to each attention head in Transformers, outperforming standard Rotary Position Embedding (RoPE) variants. The researchers show that uniform frequency schedules across heads are suboptimal, especially for long-context tasks, and that head-specific optimization improves both performance and context extension while preserving short-context accuracy.

read1 min views1 publishedJul 23, 2026

arXiv:2607.19363v1 Announce Type: new Abstract: Rotary Position Embedding (RoPE) is widely adopted in Transformers to encode positional information, yet standard implementations enforce a uniform frequency schedule and scaling across all attention heads. Using simplified retrieval tasks and length generalization scenarios, we show -- both empirically and theoretically -- that heads with different functional roles require distinct frequency ranges and attention scaling factors to operate effectively. Ignoring this structure leads to suboptimal utilization of embedding dimensions and degraded performance, particularly under long-context settings. To address these limitations, we propose AdaRoPE, which equips each attention head with learnable rotation frequencies and attention scaling factors. Pretrained LLMs with AdaRoPE consistently outperform existing RoPE variants, including partial RoPE and NoPE baselines. For context extension, we further show that uniform frequency and attention scaling, used in methods such as YaRN, are suboptimal. By applying head-specific scaling, AdaRoPE enables better context extension while better preserving short-context performance in both the extrapolation setting and the long-context continued pretraining setting. These results highlight the importance of optimizing rotary position embedding at the level of individual attention heads.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @adarope 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/adarope-not-all-atte…] indexed:0 read:1min 2026-07-23 ·