04:00
2026-09-28
arxiv.org
natural-language-processing
Manifold Projection and Iterative Autoencoder Refinement for Masked Language Modeling
A new attention-free architecture that replaces Transformer attention with a stack of autoencoder-based mixing modules achieves a significant portion of attention's performance at about 1.9x fewer FLO…