{"slug": "pretraining-latent-information-feedback-transformers-with-teacher-supervision", "title": "Pretraining Latent Information Feedback Transformers with Teacher Supervision", "summary": "Researchers introduced LIFT (Latent Information Feedback Transformer), a pretraining architecture and method that feeds deep-layer representations back to shallower layers by training models to predict both the next token and an information-dense state derived from an off-the-shelf pretrained LM's next-token distribution, according to an arXiv paper (arXiv:2609.38149v1). In experiments with pretrained models from 135M to 1B parameters, LIFT consistently outperformed standard Transformers and baselines on language modeling, downstream reasoning, and procedural tasks under a token-matched budget, and matched or beat compute-matched Transformers. A controlled state-tracking study found a tiny LIFT outperformed same-size Transformers trained on 8x more data, even when trained with states from a Transformer that fails the task.", "body_md": "arXiv:2609.38149v1 Announce Type: new \nAbstract: Transformer language models (LMs) are feed-forward: deep-layer representations are never fed back to shallower layers, and the only pathway for information to flow downward across generation steps is the decoded token. This narrow channel forces models to recompute intermediate results and to discard alternative continuations. In this work, we remove this bottleneck during pretraining, introducing the LIFT (Latent Information Feedback Transformer) architecture and training method which enable LMs to propagate state across generation. We achieve this by turning recurrent-state learning into a teacher-forced prediction problem: each input token is paired with an information-dense state, derived from the next-token distribution of an off-the-shelf pretrained LM. The model, extended with a small number of additional parameters, is then trained to predict both the next token and the next state. As the input states are precomputed, pretraining remains fully parallel across positions. At inference, the model's own predicted states are fed back, with a minor computational overhead that decreases with model size. Experiments with pretrained models ranging from 135M to 1B parameters show that LIFT consistently outperforms standard Transformers and baselines on language modeling, downstream reasoning tasks, and procedural tasks under token-matched budget, while being on par with or ahead of compute-matched Transformers. Moreover, a controlled study on a state-tracking task shows that a tiny LIFT outperforms same-size Transformers trained on 8x more data, even when trained with the states of a Transformer that fails the task. Overall, we show that LMs can learn to exploit deep-to-shallow feedback during pretraining via scalable teacher supervision.", "url": "https://wpnews.pro/news/pretraining-latent-information-feedback-transformers-with-teacher-supervision", "canonical_source": "https://www.machinebrief.com/news/pretraining-latent-information-feedback-transformers-with-te-j4tv", "published_at": "2026-09-30 04:00:00+00:00", "updated_at": "2026-09-30 05:47:09.340799+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "ai-research", "natural-language-processing"], "entities": ["LIFT", "Latent Information Feedback Transformer", "arXiv"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/pretraining-latent-information-feedback-transformers-with-teacher-supervision", "markdown": "https://wpnews.pro/news/pretraining-latent-information-feedback-transformers-with-teacher-supervision.md", "text": "https://wpnews.pro/news/pretraining-latent-information-feedback-transformers-with-teacher-supervision.txt", "jsonld": "https://wpnews.pro/news/pretraining-latent-information-feedback-transformers-with-teacher-supervision.jsonld"}}