{"slug": "manifold-projection-and-iterative-autoencoder-refinement-for-masked-language", "title": "Manifold Projection and Iterative Autoencoder Refinement for Masked Language Modeling", "summary": "A new attention-free architecture that replaces Transformer attention with a stack of autoencoder-based mixing modules achieves a significant portion of attention's performance at about 1.9x fewer FLOPs when pretrained on C4 and evaluated against parameter-matched BERT baselines, according to arXiv paper 2609.30288v1. The model uses three mixing modules — one over local neighborhoods, one over the full sequence, and one across attention heads — each compressing and reconstructing its input through a low-rank bottleneck whose width is a hyperparameter rather than a training effect. For masked positions, an iterative refinement procedure pulls each embedding toward a weighted average of its neighbors and then projects the result back to the learned manifold via an autoencoder, and a frequency-aware training schedule that samples rare tokens more than uniformly lets the model equal parameter-matched BERT and TinyBERT baselines on the rarest-token frequency bucket.", "body_md": "arXiv:2609.30288v1 Announce Type: new \nAbstract: In Transformer-based masked language models, attention is the primary mechanism for context mixing, but there are other ways to mix data across tokens. Recent attention-free mixers replace attention with fixed or hypernetwork-generated MLPs, alternating their dynamic, content-dependent weighting for computational simplicity. We build an alternative that gets the same property from a low-rank bottleneck autoencoder. We replace attention with a stack of autoencoder-based mixing modules, one operating over local neighborhoods, one over the full sequence, and one across attention heads, each compressing and reconstructing its input through a bottleneck, and its width is a hyperparameter rather than a training effect. In masked positions, we introduce an iterative refinement procedure that has two distinct steps. A pulling step that pulls an embedding representation toward a weighted average of its neighbors, and a correcting step that projects the result back to the learned manifold via an autoencoder. Our architecture achieves a significant portion of attention's performance at about $1.9 \\times$ fewer FLOPs when pretrained on C4 and evaluated with parameter-matched BERT baselines. Our model equals parameter-matched BERT and TinyBERT baselines on the rarest-token frequency bucket using a frequency-aware training schedule that samples rare tokens more than uniformly for the masking tasks.", "url": "https://wpnews.pro/news/manifold-projection-and-iterative-autoencoder-refinement-for-masked-language", "canonical_source": "https://arxiv.org/abs/2609.30288", "published_at": "2026-09-28 04:00:00+00:00", "updated_at": "2026-09-28 04:18:10.319398+00:00", "lang": "en", "topics": ["natural-language-processing", "machine-learning", "artificial-intelligence", "large-language-models", "ai-research"], "entities": ["C4", "BERT", "TinyBERT", "arXiv"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/manifold-projection-and-iterative-autoencoder-refinement-for-masked-language", "markdown": "https://wpnews.pro/news/manifold-projection-and-iterative-autoencoder-refinement-for-masked-language.md", "text": "https://wpnews.pro/news/manifold-projection-and-iterative-autoencoder-refinement-for-masked-language.txt", "jsonld": "https://wpnews.pro/news/manifold-projection-and-iterative-autoencoder-refinement-for-masked-language.jsonld"}}