Power law graph attention: exact generalization of scaled dot-product attention, empirical collapse at inference
A new arXiv paper (2608.10288v1) introduces Power Law Graph Attention (PLGA), a learned attention mechanism that exactly generalizes scaled dot-product attention (SDPA) and is used in the Power Law Decoder Representation…