cd /news/natural-language-processing/manifold-projection-and-iterative-au… · home › topics › natural-language-processing › article
[ARTICLE · art-140732] src=arxiv.org ↗ pub= topic=natural-language-processing verified=true sentiment=· neutral

Manifold Projection and Iterative Autoencoder Refinement for Masked Language Modeling

A new attention-free architecture that replaces Transformer attention with a stack of autoencoder-based mixing modules achieves a significant portion of attention's performance at about 1.9x fewer FLOPs when pretrained on C4 and evaluated against parameter-matched BERT baselines, according to arXiv paper 2609.30288v1. The model uses three mixing modules — one over local neighborhoods, one over the full sequence, and one across attention heads — each compressing and reconstructing its input through a low-rank bottleneck whose width is a hyperparameter rather than a training effect. For masked positions, an iterative refinement procedure pulls each embedding toward a weighted average of its neighbors and then projects the result back to the learned manifold via an autoencoder, and a frequency-aware training schedule that samples rare tokens more than uniformly lets the model equal parameter-matched BERT and TinyBERT baselines on the rarest-token frequency bucket.

by read1 min views1 publishedSep 28, 2026

arXiv:2609.30288v1 Announce Type: new Abstract: In Transformer-based masked language models, attention is the primary mechanism for context mixing, but there are other ways to mix data across tokens. Recent attention-free mixers replace attention with fixed or hypernetwork-generated MLPs, alternating their dynamic, content-dependent weighting for computational simplicity. We build an alternative that gets the same property from a low-rank bottleneck autoencoder. We replace attention with a stack of autoencoder-based mixing modules, one operating over local neighborhoods, one over the full sequence, and one across attention heads, each compressing and reconstructing its input through a bottleneck, and its width is a hyperparameter rather than a training effect. In masked positions, we introduce an iterative refinement procedure that has two distinct steps. A pulling step that pulls an embedding representation toward a weighted average of its neighbors, and a correcting step that projects the result back to the learned manifold via an autoencoder. Our architecture achieves a significant portion of attention's performance at about $1.9 \times$ fewer FLOPs when pretrained on C4 and evaluated with parameter-matched BERT baselines. Our model equals parameter-matched BERT and TinyBERT baselines on the rarest-token frequency bucket using a frequency-aware training schedule that samples rare tokens more than uniformly for the masking tasks.

── more in #natural-language-processing 4 stories · sorted by recency
── more on @c4 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/manifold-projection-…] indexed:0 read:1min 2026-09-28 · —