cd /news/artificial-intelligence/dual-attention-residuals · home topics artificial-intelligence article
[ARTICLE · art-67967] src=machinebrief.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

Dual Attention Residuals

Researchers propose Dual Attention Residuals (DAR), a method that brings multi-stream interaction into historical retrieval for Transformer models through reciprocal cross-stream addressing. Across dense models from 0.1B to 1B parameters and a 7B sparse-MoE model, DAR consistently improves validation loss over standard residual Transformers and Attention Residuals. The study shows that the gain cannot be explained by an additional stream or value projection alone, and that reciprocal cross-stream selection preserves depth-wise diversity.

read1 min views1 publishedJul 22, 2026

arXiv:2607.18730v1 Announce Type: new Abstract: Recent work extends Transformer residual pathways along two complementary axes: historical retrieval selects information from earlier depths, whereas multi-stream methods maintain multiple residual trajectories. These capabilities have largely been studied in isolation, and assigning an independent retriever to each stream still prevents one trajectory from influencing depth selection in another. We propose Dual Attention Residuals (DAR), which brings multi-stream interaction into historical retrieval through reciprocal cross-stream addressing. For each target stream, DAR computes depth weights from normalized states in the opposite stream and applies them to values from the target stream's own history. The retrieved states are combined for an unchanged Transformer branch and updated through constrained gated writes; a block-form variant operates on block-level histories to control overhead. Across dense models from 0.1B to 1B parameters and a 7B sparse-MoE model, DAR consistently improves validation loss over standard residual Transformers and Attention Residuals. Routing ablations show that the gain cannot be explained by an additional stream or value projection alone. Representation and intervention analyses further show that reciprocal cross-stream selection preserves depth-wise diversity and avoids the redundancy or functional imbalance observed in alternative two-stream designs.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @dual attention residuals 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/dual-attention-resid…] indexed:0 read:1min 2026-07-22 ·