cd /news/artificial-intelligence/pretraining-latent-information-feedb… · home › topics › artificial-intelligence › article
[ARTICLE · art-142308] src=machinebrief.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

Pretraining Latent Information Feedback Transformers with Teacher Supervision

Researchers introduced LIFT (Latent Information Feedback Transformer), a pretraining architecture and method that feeds deep-layer representations back to shallower layers by training models to predict both the next token and an information-dense state derived from an off-the-shelf pretrained LM's next-token distribution, according to an arXiv paper (arXiv:2609.38149v1). In experiments with pretrained models from 135M to 1B parameters, LIFT consistently outperformed standard Transformers and baselines on language modeling, downstream reasoning, and procedural tasks under a token-matched budget, and matched or beat compute-matched Transformers. A controlled state-tracking study found a tiny LIFT outperformed same-size Transformers trained on 8x more data, even when trained with states from a Transformer that fails the task.

by read1 min views1 publishedSep 30, 2026

arXiv:2609.38149v1 Announce Type: new Abstract: Transformer language models (LMs) are feed-forward: deep-layer representations are never fed back to shallower layers, and the only pathway for information to flow downward across generation steps is the decoded token. This narrow channel forces models to recompute intermediate results and to discard alternative continuations. In this work, we remove this bottleneck during pretraining, introducing the LIFT (Latent Information Feedback Transformer) architecture and training method which enable LMs to propagate state across generation. We achieve this by turning recurrent-state learning into a teacher-forced prediction problem: each input token is paired with an information-dense state, derived from the next-token distribution of an off-the-shelf pretrained LM. The model, extended with a small number of additional parameters, is then trained to predict both the next token and the next state. As the input states are precomputed, pretraining remains fully parallel across positions. At inference, the model's own predicted states are fed back, with a minor computational overhead that decreases with model size. Experiments with pretrained models ranging from 135M to 1B parameters show that LIFT consistently outperforms standard Transformers and baselines on language modeling, downstream reasoning tasks, and procedural tasks under token-matched budget, while being on par with or ahead of compute-matched Transformers. Moreover, a controlled study on a state-tracking task shows that a tiny LIFT outperforms same-size Transformers trained on 8x more data, even when trained with the states of a Transformer that fails the task. Overall, we show that LMs can learn to exploit deep-to-shallow feedback during pretraining via scalable teacher supervision.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @lift 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/pretraining-latent-i…] indexed:0 read:1min 2026-09-30 · —