Three Tokens Force Exponential Feature Rank in Nonnegative Kernel Attention
A new arXiv paper (arXiv:2608.11427v1) proves that nonnegative kernel attention requires exponentially many features—2^Ω(m)—to solve three-token Min-IP tasks with error below 1/2, whereas dense softmax attention solves t…