cd /news/artificial-intelligence/what-softmax-throws-away-mass-aware-… · home topics artificial-intelligence article
[ARTICLE · art-76339] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

What Softmax Throws Away: Mass-Aware Attention for Evidence Accumulation

A new study from arXiv identifies a fundamental limitation in standard attention mechanisms: when evidence patterns repeat, the numerator and denominator grow at the same rate, causing inputs with different amounts of accumulated evidence to produce identical aggregates. The researchers propose Mass-Aware Attention (MAA), which generalizes L1 normalization to an Lp family, allowing the numerator and denominator to scale at different rates under repetition and retaining the effective number of contributing inputs in the representation magnitude. Across four continuous-time dynamic graph models and three datasets, MAA improves future-link AUC in 11 of 12 model-dataset cells and increases linear recovery from the same hidden representation by 4.49% on average.

read1 min views1 publishedJul 28, 2026

arXiv:2607.22781v1 Announce Type: new Abstract: High task performance does not show whether a model retains prediction-relevant structural information in its internal representation. Temporal graph models, for example, can achieve high future-link AUC while basic graph statistics remain difficult to recover from the same representation. We identify one source of this gap in the weighted averaging used by standard attention: when an evidence pattern is repeated, the numerator and denominator grow at the same rate, so inputs with different amounts of accumulated evidence can produce the same aggregate. We propose Mass-Aware Attention (MAA), which generalizes standard L1 normalization to an Lp family. Under repetition, MAA makes the numerator and denominator scale at different rates, retaining the effective number of contributing inputs in the representation magnitude. It adds no supervision, parameters, hidden dimensions, or explicit count features, and recovers standard attention at p=1. Across four continuous-time dynamic graph models and three datasets, MAA improves future-link AUC in 11 of 12 model-dataset cells. Linear recovery from the same hidden representation increases by 4.49% on average, and preferential-attachment recovery improves in all 12 cells after family-wise correction. We also observe consistent evidence in marked temporal point processes, temporal knowledge graphs, retrieval-augmented generation, and spatio-temporal point processes. Information accessibility and task utility remain distinct: NLL improves in MTPP, ranking is largely preserved in TKG, additional information in RAG does not improve the diagnostic head, and downstream LayerNorm can erase the signal in STPP. These results position MAA as a general normalization principle for improving predictor-facing representation informativeness by controlling repetition invariance in standard attention.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/what-softmax-throws-…] indexed:0 read:1min 2026-07-28 ·