Manifold Projection and Iterative Autoencoder Refinement for Masked Language Modeling A new attention-free architecture that replaces Transformer attention with a stack of autoencoder-based mixing modules achieves a significant portion of attention's performance at about 1.9x fewer FLOPs when pretrained on C4 and evaluated against parameter-matched BERT baselines, according to arXiv paper 2609.30288v1. The model uses three mixing modules — one over local neighborhoods, one over the full sequence, and one across attention heads — each compressing and reconstructing its input through a low-rank bottleneck whose width is a hyperparameter rather than a training effect. For masked positions, an iterative refinement procedure pulls each embedding toward a weighted average of its neighbors and then projects the result back to the learned manifold via an autoencoder, and a frequency-aware training schedule that samples rare tokens more than uniformly lets the model equal parameter-matched BERT and TinyBERT baselines on the rarest-token frequency bucket. arXiv:2609.30288v1 Announce Type: new Abstract: In Transformer-based masked language models, attention is the primary mechanism for context mixing, but there are other ways to mix data across tokens. Recent attention-free mixers replace attention with fixed or hypernetwork-generated MLPs, alternating their dynamic, content-dependent weighting for computational simplicity. We build an alternative that gets the same property from a low-rank bottleneck autoencoder. We replace attention with a stack of autoencoder-based mixing modules, one operating over local neighborhoods, one over the full sequence, and one across attention heads, each compressing and reconstructing its input through a bottleneck, and its width is a hyperparameter rather than a training effect. In masked positions, we introduce an iterative refinement procedure that has two distinct steps. A pulling step that pulls an embedding representation toward a weighted average of its neighbors, and a correcting step that projects the result back to the learned manifold via an autoencoder. Our architecture achieves a significant portion of attention's performance at about $1.9 \times$ fewer FLOPs when pretrained on C4 and evaluated with parameter-matched BERT baselines. Our model equals parameter-matched BERT and TinyBERT baselines on the rarest-token frequency bucket using a frequency-aware training schedule that samples rare tokens more than uniformly for the masking tasks.