cd /news/machine-learning/dust-pretraining-transformers-withou… · home › topics › machine-learning › article
[ARTICLE · art-145675] src=qlabs.sh ↗ pub= topic=machine-learning verified=true sentiment=↑ positive

Dust: Pretraining Transformers Without Backpropagation

Researchers Samip Dahal, Bishwas Mandal, Serdar Gülbahar and Akshay Vegesna released Dust, a zeroth-order pretraining method that perturbs activations independently at every token so one forward pass evaluates a virtual population in parallel, claiming it is the first such method competitive with backpropagation for pretraining transformer language models. The authors report Dust is roughly 10^3 to 10^4 times more efficient than EGGROLL, a state-of-the-art evolution-strategies method, from 1M tokens up based on their extrapolations, and that a 243M-parameter model outperforms a 120x smaller model at most population sizes, with gradient estimates staying aligned with backprop up to 1B tokens.

read28 min views1 publishedOct 5, 2026
@misc{dahal2026backprop,
  title  = {Dust: Pretraining Transformers Without Backpropagation},
  author = {Dahal, Samip and Mandal, Bishwas and G{\"u}lbahar, Serdar and Vegesna, Akshay},
  year   = {2026},
  url    = {https://qlabs.sh/research/dust}
}

TL;DR #

  • We present the first zeroth-order method that is competitive with backprop at pretraining transformer language models. Dust perturbs activations (node perturbation) independently at every token, so each token is a virtual population member and one forward pass evaluates them all in parallel.
  • Dust approximates backprop closely at large population (i.e. substantially more compute) and in multiple settings even exceeds it. This hints that in a compute-rich regime we might be able to surpass backprop.
  • Dust is orders of magnitude more efficient than weight-space ES. From 1M tokens up, Dust is on the order of $10^3$ to $10^4$ times more efficient than a transformer implementation of EGGROLL, a state-of-the-art ES method, based on our extrapolations.
  • Zeroth-order methods are widely believed not to scale to large networks. Strikingly, we find larger models are more population-efficient, not less: a 243M-parameter model outperforms a $120\times$ smaller model at most population sizes.
  • Dust’s gradient estimates align better with backprop’s as population grows, and stay well aligned at every scale we test, up to 1B tokens, which is encouraging for scaling.

1 Introduction #

Deep learning has been built around backprop, the only credit assignment algorithm capable of training modern neural nets, including transformer-based language models. Backprop requires differentiability and produces first-order gradients, and deep learning’s architectures, optimizers, and hardware have co-evolved around this constraint.

However, as the amount of compute available in the world increases, we might prefer more generic and brute-force learning algorithms based on search over inductive biases like differentiability, backprop, and approximations of higher-order gradients. The bitter lesson (Sutton, 2019Richard S. Sutton. The bitter lesson. http://www.incompleteideas.net/IncIdeas/BitterLesson.html, 2019. Blog post.) is that general methods that scale with compute eventually win, and AlphaGo Zero (Silver et al., 2017David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, Yutian Chen, Timothy Lillicrap, Fan Hui, Laurent Sifre, George van den Driessche, Thore Graepel, and Demis Hassabis. Mastering the game of Go without human knowledge. Nature, 550 (7676): 354–359, 2017. doi: 10.1038/nature24270.) is the obvious example. Bootstrapping AlphaGo on human data helped the network learn faster initially, but with a lot of computation the purely self-play network overtook it. Similarly, differentiability and backprop might be good inductive biases in the low-compute regime, where they make learning efficient, but in the high-compute regime they limit the space of architectures that work. Even within an architecture, gradient-based methods fail to explore the loss landscape optimally (Liu et al., 2020Shengchao Liu, Dimitris Papailiopoulos, and Dimitris Achlioptas. Bad global minima exist and SGD can reach them. In Advances in Neural Information Processing Systems, volume 33, 2020.). This might also explain why current neural nets require massive amounts of data to generalize. A more flexible credit assignment algorithm based on search is likely an important step towards much better generalization.

In this paper, we aim to replace backprop with a learning algorithm based much more on brute-force computation and much less on analytic structure. We call it Dust. Dust is a zeroth-order optimization algorithm that perturbs activations, rewards each perturbation by how much it lowers the loss, and averages the reward-weighted perturbations over a population to estimate the gradient. Traditional ES methods that perturb weights (Salimans et al., 2017Tim Salimans, Jonathan Ho, Xi Chen, Szymon Sidor, and Ilya Sutskever. Evolution strategies as a scalable alternative to reinforcement learning. arXiv preprint arXiv:1703.03864, 2017.), like EGGROLL (Sarkar et al., 2025Bidipta Sarkar, Mattie Fellows, Juan Agustin Duque, Alistair Letcher, Antonio León Villares, Anya Sims, Clarisse Wibault, Dmitry Samsonov, Dylan Cope, Jarek Liesen, Kang Li, Lukas Seier, Theo Wolf, Uljad Berdica, Valentin Mohl, Alexander David Goldie, Aaron Courville, Karin Sevegnani, Shimon Whiteson, and Jakob Nicolaus Foerster. Evolution strategies at the hyperscale. arXiv preprint arXiv:2511.16652, 2025.), scale with population, but scaling the population is costly because each member must be materialized and evaluated. We remove both costs with the concept of virtual population, where we avoid materializing every member by bypassing weight space entirely and instead perturb activations, as in node perturbation (Werfel et al., 2003Justin Werfel, Xiaohui Xie, and H. Sebastian Seung. Learning curves for stochastic gradient descent in linear feedforward networks. In Advances in Neural Information Processing Systems, volume 16, 2003.; Widrow and Lehr, 1990Bernard Widrow and Michael A. Lehr. 30 years of adaptive neural networks: Perceptron, Madaline, and backpropagation. Proceedings of the IEEE, 78 (9): 1415–1442, 1990. doi: 10.1109/5.58323.). We do so independently at every token, so each token is a member and one forward pass evaluates them all in parallel.

Activations are a more interesting space to search over than weights. Mechanistic interpretability has shown that reasoning, whether verbalizable or not, lives in the activations (Gurnee et al., 2026Wes Gurnee, Nicholas Sofroniew, Adam Pearce, Mateusz Piotrowski, Isaac Kauvar, Runjin Chen, Anna Soligo, Paul Bogdan, Euan Ong, Rowan Wang, Ben Thompson, David Abrahams, Subhash Kantamneni, Emmanuel Ameisen, Joshua Batson, and Jack Lindsey. Verbalizable representations form a global workspace in language models. arXiv preprint arXiv:2607.15495, 2026.; Lindsey et al., 2025Jack Lindsey, Wes Gurnee, Emmanuel Ameisen, Brian Chen, Adam Pearce, Nicholas L. Turner, Craig Citro, et al. On the biology of a large language model. Transformer Circuits Thread, 2025.), which means this approach could turn training into a search over latent reasoning (Vegesna and Dahal, 2025Akshay Vegesna and Samip Dahal. Decoupling search and learning in neural net training. arXiv preprint arXiv:2509.10973, 2025.). We then pair the activation-space perturbation with a very generic credit assignment rule that assigns different token-level rewards to different layer types in a transformer block. Those two biases, along with a few implementation details and efficiency measures, like avoiding interference between perturbed modules, are the whole algorithm.

We make the following contributions.

  • We present the first zeroth-order method that is competitive with backprop at pretraining transformer language models. At large populations Dust exceeds backprop in multiple settings, which suggests that in a compute-rich regime we might be able to surpass backprop.
  • Dust is orders of magnitude more efficient than weight-space ES. From 1M tokens up, Dust is on the order of $10^3$ to $10^4$ times more efficient than a transformer implementation of EGGROLL, based on our extrapolations.
  • Contrary to conventional wisdom, larger models are often more population-efficient, not less, and can make use of larger populations. This gives a new view of overparameterization as a larger search space with potentially better geometry.
  • Dust’s gradient estimates align better with backprop’s as the population grows, and the alignment holds up at every scale we test, up to 1B tokens, which is encouraging for scaling.

The goal of this paper is to lay the foundations of a search-based credit assignment algorithm that is competitive with backprop on the hardest task we could think of: pretraining transformers. We do not attempt to make it compute-efficient enough to replace backprop today. We also do not train the new kinds of neural nets it makes accessible, like nets with an external program in the loop or transformers looped over many steps that backpropagation through time struggles to train. Both are left to future work.

2 Method #

Dust works as follows. We add Gaussian noise to the output of each linear layer, independently at every token, run a forward pass, and reward each token’s noise by the change in loss at that token. The reward-weighted noise, averaged over draws, is the estimated error at the layer’s output, and its outer product with the layer’s input is the weight gradient. Attention internals get a variant of it: they are credited through the estimated error at the attention output over current and future tokens, instead of the tokens’ loss directly. The core intuition is that while weight-space ES evaluates one population member per forward pass, we evaluate one per token, in parallel, and a member is materialized by adding noise to a hidden state, which is cheap. On a modern transformer a single forward pass therefore evaluates a population at least three orders of magnitude larger than weight-space ES. We describe each component in detail below.

2.1 Activation-Space Perturbation

The bottleneck of evolution strategies is population size. Every member needs its own perturbed copy of the weights and its own forward pass. EGGROLL (Sarkar et al., 2025Bidipta Sarkar, Mattie Fellows, Juan Agustin Duque, Alistair Letcher, Antonio León Villares, Anya Sims, Clarisse Wibault, Dmitry Samsonov, Dylan Cope, Jarek Liesen, Kang Li, Lukas Seier, Theo Wolf, Uljad Berdica, Valentin Mohl, Alexander David Goldie, Aaron Courville, Karin Sevegnani, Shimon Whiteson, and Jakob Nicolaus Foerster. Evolution strategies at the hyperscale. arXiv preprint arXiv:2511.16652, 2025.) makes the copies cheap with low-rank perturbations, but each member is still one sequence element of the batch, so the population is bounded by the forward passes one can afford. We perturb activations instead, independently at every token. At that token the network behaves as if a low-rank perturbation had been applied to the weights of the layer that produced the activations, without the perturbation ever being materialized in the weights. We call this a virtual population. A sequence in a transformer has a few thousand tokens, so one forward pass evaluates a few thousand members per sequence instead of one. Every weight in the model is trained this way except the $2L$ residual mixing scalars, which are trained by ordinary weight-space ES.

Adding noise to activations rather than weights is node perturbation (Widrow and Lehr, 1990Bernard Widrow and Michael A. Lehr. 30 years of adaptive neural networks: Perceptron, Madaline, and backpropagation. Proceedings of the IEEE, 78 (9): 1415–1442, 1990. doi: 10.1109/5.58323.), and the usual argument for it is dimensionality (Ren et al., 2023Mengye Ren, Simon Kornblith, Renjie Liao, and Geoffrey Hinton. Scaling forward gradient with local losses. In International Conference on Learning Representations, 2023.; Werfel et al., 2003Justin Werfel, Xiaohui Xie, and H. Sebastian Seung. Learning curves for stochastic gradient descent in linear feedforward networks. In Advances in Neural Information Processing Systems, volume 16, 2003.). A layer’s output has $d_{\mathrm{out}}$ entries and its weights have $d_{\mathrm{out}} \times d_{\mathrm{in}}$, so activation noise lives in a much smaller space. Naively, the dimensionality argument doesn’t hold for transformers with large activations across many tokens. The noise on one sequence is a $T \times d_{\mathrm{out}}$ tensor, which has at least as many entries as the weight matrix once $T \ge d_{\mathrm{in}}$. However, with per-token independent perturbations and rewards, what perturbing activations gives is a new, efficient population along the token axis that is orthogonal to the batch axis EGGROLL already relies on.

2.2 Credit Assignment

For a linear layer $y_t = W x_t$ we jitter its output at all tokens, $y_t \to y_t + \sigma a_t$ with $a_t \sim \mathcal{N}(0, I)$ and $\sigma$ the noise scale, and run the forward pass. At each token $s$ we calculate the centered loss reduction $c_s = \tilde{\ell}_s - \ell_s$, where $\ell_s$ is the perturbed loss and $\tilde{\ell}_s$ is the mean perturbed loss at that token across draws evaluated together in the same batched forward. The reward of the jitter at token $t$ is the loss reduction at $t$ and, with a decay $\gamma$, the loss reductions at the tokens after it, which the jitter also reaches through attention,

With $\gamma = 0$ a jitter is rewarded by its own token alone. We leave it to the tuning of Section 2.3 to decide which layers see the future tokens. One independent jitter of all tokens is a draw, and a population is $K$ draws. Averaged over the draws, the reward-weighted noise

is the estimated error at the layer’s output, and its outer product with the layer’s input, which the forward pass already computed, summed over tokens, is the weight gradient,

Backprop forms the same outer product with the same input. The only difference is that it gets the output error from the chain rule and we get it from the population. For an embedding layer $x_t$ is one-hot, so the outer product is a scatter-add of $\hat g_t$ into the token’s row.

With unlimited population the estimator above is the whole method, and every perturbation could go into one forward pass. Each draw’s reward-weighted noise is the gradient plus an error with no preferred direction. Over draws the gradient adds up linearly while the errors add up as the square root, so their ratio falls as the population grows and the interference vanishes in the limit.

2.3 Interference and Tuning

At a population we can afford, the main cost is interference, i.e. we jitter many layers at many tokens in the same forward pass, so the loss change that rewards one token’s noise also picks up the effect of every other perturbation in that pass. We reduce such interference in three ways. First, different layer types are jittered in separate forward passes, each with its own noise scale, and each block gets its own passes. These passes are cheaper than full forwards because the clean forward is cached and a draw for block $l$ only reruns blocks $l$ onward. Second, the attention internals (query, key, value, gate, value embedding) are jittered separately. Token losses barely register their jitters, so they are rewarded through the attention output instead, as described below. Third, the language modeling head is jittered directly on the cached logits, re-evaluating only the cross-entropy and only a slab of the vocabulary per draw, which costs a small fraction of a forward pass and lets the head run a much larger population.

For the attention internals we recompute only their block’s attention output with the jitter, from the cached clean activations. We score the jitters using the alignment with the estimated attention output gradients,

where $\Delta o_s$ is the change the jitter makes in the attention output at token $s$ and $\hat g_s$ is that output’s estimated error from Equation 2. The reward is Equation 1 with these scores, computed per head, in place of the loss reductions.

The hyperparameters, the noise scale of each layer type, the credit decay of the attention internals and the share of the population each layer gets, can be tuned in two ways. One is a grid search that trains with each setting on a small token budget and keeps the setting that lowers the loss most. It is reliable but expensive. The other is a grid search that maximizes the cosine between our estimate and the backprop gradient on a single batch, which needs no training at all. A larger cosine on one batch doesn’t always lower the loss after training, so the cosine picks candidates and training decides. Either way the tuning is mostly a one-time cost, since the settings it finds largely generalize across token budgets and population sizes, with one exception. At the largest population at 10M and 20M tokens, a slower credit decay and a shift of draws toward the attention side still pay (Appendix F). Hence, this search recovers general principles of the method rather than settings for one run. As expected, all layers require no credits from future tokens except the keys, values, gates and value embeddings, which are read by the later tokens that attend to them and get a $\gamma$ close to one.

3 Pretraining Without a Backward Pass #

3.1 Setup

We train GPT-style transformers on FineWeb with a 4096-token BPE tokenizer, a batch of 16k tokens (8 sequences of 2048 tokens), one epoch, and SGD with momentum at a constant learning rate. The base model has 8 layers and width 512. Every method gets the same protocol and three seeds per cell, and is tuned separately at every token budget and population. Dust and backprop share one momentum and learning-rate grid; we implement EGGROLL with the same transformer architecture (named EGGROLL-Transformer) and it is tuned over its own grid of step size, momentum, noise scale and fitness shaping. Validation and test are held-out sets of 544 sequences each. We report the test loss at best validation checkpoint.

We count population in draws for Dust and in forward passes of the batch for EGGROLL (Sarkar et al., 2025Bidipta Sarkar, Mattie Fellows, Juan Agustin Duque, Alistair Letcher, Antonio León Villares, Anya Sims, Clarisse Wibault, Dmitry Samsonov, Dylan Cope, Jarek Liesen, Kang Li, Lukas Seier, Theo Wolf, Uljad Berdica, Valentin Mohl, Alexander David Goldie, Aaron Courville, Karin Sevegnani, Shimon Whiteson, and Jakob Nicolaus Foerster. Evolution strategies at the hyperscale. arXiv preprint arXiv:2511.16652, 2025.), our weight-space baseline. A draw is one jitter of activations across all tokens of selected layers, rewarded with the token losses, and the population $K$ is the number of draws per update. A draw is slightly cheaper than a forward pass, because the clean forward is cached and a draw that jitters block $l$ reruns only the blocks from $l$ on. $K$ leaves out the draws of the head and the attention internals, which are a small fraction of the update’s FLOPs (Appendix E). Taken together, from a population of 256 up Dust uses less compute than EGGROLL at the same population, so the comparison is lenient to EGGROLL.

3.2 Main Results

We sweep the token budget from 100k to 20M against populations from 64 to 16k, with backprop tuned on the same grid at every budget (Table 1, Figure 2). At 100k and 1M tokens Dust ends below backprop, from a few hundred draws at 100k and from a thousand at 1M. At 10M and 20M tokens, the gap with backprop shrinks with population. At 10M the ladder has flattened and its fitted limit lands just above backprop. At 20M the ladder is still falling at 16k draws. Its power law fit puts the limit at 4.431 (95% interval 3.89 to 4.58), below backprop’s 4.633, but with the ladder still falling the fit is loosely constrained, so we read it as evidence that the gap keeps closing with population rather than as a measured limit.

Weight-space ES is far less efficient. With 256 times the population, EGGROLL at 16k still does not reach Dust at 64 draws. It comes within 0.02 at 100k tokens and stays 0.4 to 0.6 above at 1M, 10M and 20M (Figure 2). EGGROLL’s ladder is still falling steeply at 16k, so it would keep improving with more population, but continuing the ladder puts what it needs to match Dust’s smallest population at several thousand to about $10^4$ times that population (Appendix D). Figure 1 also shows the gap during training; Dust at 1k and 16k draws stays close to backprop’s curve throughout, while every EGGROLL curve falls behind early and flattens well above Dust at 64 draws.

3.3 Dust Under Adam

While we focus primarily on SGD for the rest of the paper, we repeat the 1M token ladder with Adam on all three methods (Figure 3), re-tuning backprop’s learning rates, Dust’s own hyperparameters and EGGROLL’s step size, momentum and fitness shaping at every population. Interestingly, EGGROLL gains almost nothing from Adam and its tuned Adam ladder lands within 0.01 of its SGD ladder in Table 1 at every population. However, Adam improves both Dust and backprop and leaves the ladder’s shape similar to before. Dust closes on backprop with a large population and the interval on its limit sits below backprop (Figure 3). So Dust’s estimate already works with modern optimizers, even though modern optimizers were optimized for backprop gradients. We suspect that coevolution of optimizers with Dust could lead to further gains and leave this to future work.

4 Search in High-Dimensional Space #

4.1 Overparameterization

Population Parameters Across sizes
2.0M 7.3M 38M 243M
64 5.705 5.556 5.558 5.719
256 5.486 5.362 5.358 5.419
1k 5.265 5.161 5.158 5.214
4k 5.189 5.095 5.053 5.124
16k 5.171 5.065 5.036 5.086
Backprop 5.180 5.066 5.015 5.048

The conventional view is that zeroth-order methods cannot train large networks (Lillicrap et al., 2020Timothy P. Lillicrap, Adam Santoro, Luke Marris, Colin J. Akerman, and Geoffrey Hinton. Backpropagation and the brain. Nature Reviews Neuroscience, 21 (6): 335–346, 2020. doi: 10.1038/s41583-020-0277-3.). A forward pass returns a single scalar, so the variance of the gradient estimate grows with the number of perturbed dimensions, and with it the population needed for a useful update (Nesterov and Spokoiny, 2017Yurii Nesterov and Vladimir Spokoiny. Random gradient-free minimization of convex functions. Foundations of Computational Mathematics, 17 (2): 527–566, 2017. doi: 10.1007/s10208-015-9296-2.; Werfel et al., 2003Justin Werfel, Xiaohui Xie, and H. Sebastian Seung. Learning curves for stochastic gradient descent in linear feedforward networks. In Advances in Neural Information Processing Systems, volume 16, 2003.). We test this directly by training four sizes, 2M, 7M, 38M and 243M parameters (a $120\times$ range in parameter count), at a fixed 10M tokens across the population sweep (Figure 4; see Appendix F for tuning details). We make two striking observations that challenge conventional wisdom.

  • Bigger models are often more population-efficient, not less. At every population from 256 up the loss falls from 2M to 7M to 38M parameters, and the 243M model is only slightly worse. Even at the smallest population we test, a model $120\times$ larger does about the same, and is better at every other population. Up to a point, a larger model gains more from its size than it loses to variance. What we observe, however, is that the gap to backprop at the same population and model size increases with model size, but barely noticeable, esp at large population sizes.
  • Bigger models keep improving with population where small ones flatten. Small models saturate sooner, while big models keep improving with large population sizes. Past 1k draws the 38M and 243M models gain about 30% more than the 2M and 7M ones, and no amount of population makes the tiny model competitive.

The right way to think about model size is therefore as the size and the geometry of the search space. A larger model has a larger space to search over, which is what lets it put a large population to use, and potentially a better-conditioned loss landscape geometry, which could be why the search is more effective even at small populations.

4.2 Emergence of Backprop-Like Gradients

We measure the cosine between Dust’s estimate and the backprop gradient on the same batch, per layer type and per layer, on backprop-trained checkpoints spanning two orders of magnitude in tokens, i.e. from 10M to 1B tokens, with the population growing from 64 to 128k forward passes per step (Figure 5). The cosine rises with population for every layer type at every stage of training, and a two-parameter law

fits each layer type with RMSE below 0.06, with $c_{\max}$ the ceiling and $c$ the population at which the layer reaches $c_{\max}/\sqrt{2}$. Useful gradients emerge from the large population alone, with nothing about the chain rule built in, just light tuning of the hyperparameters to maximize the cosine similarity as described in Section 2.3. For the 100M token checkpoint, the fits to layer type mean cosine similarities have RMSEs of $0.0032$–$0.0367$ (see Appendix G for details). Dust’s gradients approach backprop’s in cosine, but how close they get varies by layer type and by layer. EGGROLL’s estimate, measured the same way at the same number of forward passes, is growing with population but is still below 0.05 in cosine at 128k for every type but the head, which might explain why training on it went nowhere in Section 3.2.

Importantly, the cosines hold up across most layers as the token count grows, which is encouraging for scaling. Section 3 shows the population required to match and exceed backprop grows with tokens. However, the cosine stays flat across two orders of magnitude in tokens at large populations, which means that at a large enough population it’s possible that this requirement no longer grows with more tokens. Furthermore, the fact that Dust’s gradients approach backprop’s but don’t converge to them exactly is in fact a good property. The estimate points in a similar direction without being backprop’s gradient, which leads to a different optimization trajectory, and in our experiments that trajectory can even be better than backprop’s (Section 3.2).

5 Conclusion #

Since Rumelhart et al. (1986)David E. Rumelhart, Geoffrey E. Hinton, and Ronald J. Williams. Learning representations by back-propagating errors. Nature, 323: 533–536, 1986., backprop has been the algorithm that trains neural nets, and the architectures, optimizers and hardware we have were all built around it. As compute becomes more abundant, we think much better alternatives are possible. We introduced Dust, an algorithm that drastically improves upon existing ES algorithms and approximates backprop closely at pretraining transformers, even exceeding it with large amounts of computation.

There are many interesting open questions. The first is whether, and how, Dust can find better directions than backprop’s first-order gradient by implicitly exploring the loss landscape, picking up higher-order curvature that pulls the search toward flat regions. We have hints that it can, since at large populations it sometimes ends up below backprop, but the mechanism is not clear. The second is that Dust opens up the search space over architectures, since it does not need the network to be end-to-end differentiable, and it may do better where backprop is known to struggle, like recurrent or looped computation trained by backpropagation through time. The third is compute efficiency, which was not the focus of this paper. We must have orders of magnitude more compute efficiency before Dust becomes a practical alternative to backprop at current levels of compute.

Early work on zeroth-order pretraining of language models (Allaire et al., 2025Nathan Allaire, Mahsa Ghazvini Nejad, Sébastien Le Digabel, and Vahid Partovi Nia. Zeroth order optimization for pretraining language models. In Proceedings of ICPRAM, pages 113–121, 2025. doi: 10.5220/0013261100003905. URL https://doi.org/10.5220/0013261100003905.) examined the difficulty of training transformers from scratch with weight perturbations. Their later method, KronZO (Allaire et al., 2026Nathan Allaire, Sébastien Le Digabel, Dominique Orban, and Vahid Partovi Nia. Zeroth-order Kronecker optimization for pretraining language models. SN Computer Science, 7, 2026. URL https://www.gerad.ca/en/papers/G-2025-44. Article 162.), uses compact perturbations with Kronecker structure and selective directional updates to improve pretraining while reducing memory use. EGGROLL (Sarkar et al., 2025Bidipta Sarkar, Mattie Fellows, Juan Agustin Duque, Alistair Letcher, Antonio León Villares, Anya Sims, Clarisse Wibault, Dmitry Samsonov, Dylan Cope, Jarek Liesen, Kang Li, Lukas Seier, Theo Wolf, Uljad Berdica, Valentin Mohl, Alexander David Goldie, Aaron Courville, Karin Sevegnani, Shimon Whiteson, and Jakob Nicolaus Foerster. Evolution strategies at the hyperscale. arXiv preprint arXiv:2511.16652, 2025.) makes large populations of weight perturbations efficient on GPUs through low-rank structure. Dust instead searches over activations and uses rewards for each token to extract more credit from each forward pass.

For fine tuning, MeZO (Malladi et al., 2023Sadhika Malladi, Tianyu Gao, Eshaan Nichani, Alex Damian, Jason D. Lee, Danqi Chen, and Sanjeev Arora. Fine-tuning language models with just forward passes. In Advances in Neural Information Processing Systems, volume 36, 2023.) showed that language models can be adapted with forward passes alone and memory use close to inference. Evolution Strategies at Scale (Qiu et al., 2026Xin Qiu, Yulu Gan, Conor F. Hayes, Qiyao Liang, Yinggan Xu, Roberto Dailey, Elliot Meyerson, Babak Hodjat, and Risto Miikkulainen. Evolution strategies at scale: LLM fine-tuning beyond reinforcement learning. In Proceedings of the International Conference on Machine Learning, 2026. URL https://arxiv.org/abs/2509.24372. arXiv:2509.24372.) demonstrates fine tuning of all parameters with ES in language models with billions of parameters. Neural Thickets (Gan and Isola, 2026Yulu Gan and Phillip Isola. Neural thickets: Diverse task experts are dense around pretrained weights. arXiv preprint arXiv:2603.12228, 2026. URL https://arxiv.org/abs/2603.12228.) finds useful task experts by randomly perturbing pretrained weights, selecting the best candidates and ensembling their predictions. These results show how much search can achieve around a pretrained model. Our experiments address learning the representations themselves through pretraining from scratch.

Our activation perturbations build on node perturbation (Werfel et al., 2003Justin Werfel, Xiaohui Xie, and H. Sebastian Seung. Learning curves for stochastic gradient descent in linear feedforward networks. In Advances in Neural Information Processing Systems, volume 16, 2003.). GEMINI (Le Cun et al., 1988Yann Le Cun, Conrad C. Galland, and Geoffrey E. Hinton. GEMINI: Gradient estimation through matrix inversion after noise injection. In Advances in Neural Information Processing Systems, volume 1, pages 141–148, 1988. URL https://papers.neurips.cc/paper_files/paper/1988/file/a0a080f42e6f13b3a2df133f073095dd-Paper.pdf.) injected noise into the first hidden layer and recovered layerwise gradient estimates through iterative matrix inversion. Zoop (Hu et al., 2025Xixi Hu, Bo Liu, Qiang Liu, Xiaocong Du, Bhargav Bhushanam, Louis Feng, Chengyue Gong, and Kaizhao Liang. Zoop it! Efficient zero-order optimization with output perturbation. In ICML Workshop on Tiny Titans: The next wave of On-Device Learning for Foundation Models, 2025. URL https://openreview.net/forum?id=Tc8vFyRhPO.) uses output perturbations for language model fine tuning, converting estimated output gradients into parameter updates with local derivatives. Scaling Forward Gradient with Local Losses (Ren et al., 2023Mengye Ren, Simon Kornblith, Renjie Liao, and Geoffrey Hinton. Scaling forward gradient with local losses. In International Conference on Learning Representations, 2023.) combines activation perturbations and local losses with forward mode automatic differentiation to reduce estimator variance. Forward gradients with multiple tangents (Flügel et al., 2025Katharina Flügel, Daniel Coquelin, Marie Weiel, Charlotte Debus, Achim Streit, and Markus Götz. Beyond backpropagation: Optimization with multi-tangent forward gradients. In International Joint Conference on Neural Networks, 2025. URL https://arxiv.org/pdf/2410.17764v2.) also use forward mode differentiation, combining multiple directional derivatives through orthogonal projection to improve gradient estimates.

A separate line of work replaces the global backward pass with local learning dynamics: Sakana AI’s PC-ALM (Seely and Gould, 2026Jeffrey Seely and Julian Gould. Augmented Lagrangian predictive coding. arXiv preprint arXiv:2605.31022, 2026. URL https://arxiv.org/abs/2605.31022.) propagates credit through local predictive coding dynamics and Lagrange multipliers. It uses local derivatives and is evaluated on image classification tasks. Dust estimates credit from forward perturbations during transformer pretraining, combining rewards for each token with local targets for attention outputs.

References #

Nathan Allaire, Mahsa Ghazvini Nejad, Sébastien Le Digabel, and Vahid Partovi Nia. Zeroth order optimization for pretraining language models. In Proceedings of ICPRAM, pages 113–121, 2025. doi: 10.5220/0013261100003905. URL https://doi.org/10.5220/0013261100003905.

Nathan Allaire, Sébastien Le Digabel, Dominique Orban, and Vahid Partovi Nia. Zeroth-order Kronecker optimization for pretraining language models. SN Computer Science, 7, 2026. URL https://www.gerad.ca/en/papers/G-2025-44. Article 162.

Katharina Flügel, Daniel Coquelin, Marie Weiel, Charlotte Debus, Achim Streit, and Markus Götz. Beyond backpropagation: Optimization with multi-tangent forward gradients. In International Joint Conference on Neural Networks, 2025. URL https://arxiv.org/pdf/2410.17764v2.

Yulu Gan and Phillip Isola. Neural thickets: Diverse task experts are dense around pretrained weights. arXiv preprint arXiv:2603.12228, 2026. URL https://arxiv.org/abs/2603.12228.

Wes Gurnee, Nicholas Sofroniew, Adam Pearce, Mateusz Piotrowski, Isaac Kauvar, Runjin Chen, Anna Soligo, Paul Bogdan, Euan Ong, Rowan Wang, Ben Thompson, David Abrahams, Subhash Kantamneni, Emmanuel Ameisen, Joshua Batson, and Jack Lindsey. Verbalizable representations form a global workspace in language models. arXiv preprint arXiv:2607.15495, 2026.

Xixi Hu, Bo Liu, Qiang Liu, Xiaocong Du, Bhargav Bhushanam, Louis Feng, Chengyue Gong, and Kaizhao Liang. Zoop it! Efficient zero-order optimization with output perturbation. In ICML Workshop on Tiny Titans: The next wave of On-Device Learning for Foundation Models, 2025. URL https://openreview.net/forum?id=Tc8vFyRhPO.

Yann Le Cun, Conrad C. Galland, and Geoffrey E. Hinton. GEMINI: Gradient estimation through matrix inversion after noise injection. In Advances in Neural Information Processing Systems, volume 1, pages 141–148, 1988. URL https://papers.neurips.cc/paper_files/paper/1988/file/a0a080f42e6f13b3a2df133f073095dd-Paper.pdf.

Timothy P. Lillicrap, Adam Santoro, Luke Marris, Colin J. Akerman, and Geoffrey Hinton. Backpropagation and the brain. Nature Reviews Neuroscience, 21 (6): 335–346, 2020. doi: 10.1038/s41583-020-0277-3.

Jack Lindsey, Wes Gurnee, Emmanuel Ameisen, Brian Chen, Adam Pearce, Nicholas L. Turner, Craig Citro, et al. On the biology of a large language model. Transformer Circuits Thread, 2025.

Shengchao Liu, Dimitris Papailiopoulos, and Dimitris Achlioptas. Bad global minima exist and SGD can reach them. In Advances in Neural Information Processing Systems, volume 33, 2020.

Sadhika Malladi, Tianyu Gao, Eshaan Nichani, Alex Damian, Jason D. Lee, Danqi Chen, and Sanjeev Arora. Fine-tuning language models with just forward passes. In Advances in Neural Information Processing Systems, volume 36, 2023.

Yurii Nesterov and Vladimir Spokoiny. Random gradient-free minimization of convex functions. Foundations of Computational Mathematics, 17 (2): 527–566, 2017. doi: 10.1007/s10208-015-9296-2.

Xin Qiu, Yulu Gan, Conor F. Hayes, Qiyao Liang, Yinggan Xu, Roberto Dailey, Elliot Meyerson, Babak Hodjat, and Risto Miikkulainen. Evolution strategies at scale: LLM fine-tuning beyond reinforcement learning. In Proceedings of the International Conference on Machine Learning, 2026. URL https://arxiv.org/abs/2509.24372. arXiv:2509.24372.

Mengye Ren, Simon Kornblith, Renjie Liao, and Geoffrey Hinton. Scaling forward gradient with local losses. In International Conference on Learning Representations, 2023.

David E. Rumelhart, Geoffrey E. Hinton, and Ronald J. Williams. Learning representations by back-propagating errors. Nature, 323: 533–536, 1986.

Tim Salimans, Jonathan Ho, Xi Chen, Szymon Sidor, and Ilya Sutskever. Evolution strategies as a scalable alternative to reinforcement learning. arXiv preprint arXiv:1703.03864, 2017.

Bidipta Sarkar, Mattie Fellows, Juan Agustin Duque, Alistair Letcher, Antonio León Villares, Anya Sims, Clarisse Wibault, Dmitry Samsonov, Dylan Cope, Jarek Liesen, Kang Li, Lukas Seier, Theo Wolf, Uljad Berdica, Valentin Mohl, Alexander David Goldie, Aaron Courville, Karin Sevegnani, Shimon Whiteson, and Jakob Nicolaus Foerster. Evolution strategies at the hyperscale. arXiv preprint arXiv:2511.16652, 2025.

Jeffrey Seely and Julian Gould. Augmented Lagrangian predictive coding. arXiv preprint arXiv:2605.31022, 2026. URL https://arxiv.org/abs/2605.31022.

David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, Yutian Chen, Timothy Lillicrap, Fan Hui, Laurent Sifre, George van den Driessche, Thore Graepel, and Demis Hassabis. Mastering the game of Go without human knowledge. Nature, 550 (7676): 354–359, 2017. doi: 10.1038/nature24270.

Richard S. Sutton. The bitter lesson. http://www.incompleteideas.net/IncIdeas/BitterLesson.html, 2019. Blog post.

Akshay Vegesna and Samip Dahal. Decoupling search and learning in neural net training. arXiv preprint arXiv:2509.10973, 2025.

Justin Werfel, Xiaohui Xie, and H. Sebastian Seung. Learning curves for stochastic gradient descent in linear feedforward networks. In Advances in Neural Information Processing Systems, volume 16, 2003.

Bernard Widrow and Michael A. Lehr. 30 years of adaptive neural networks: Perceptron, Madaline, and backpropagation. Proceedings of the IEEE, 78 (9): 1415–1442, 1990. doi: 10.1109/5.58323.

── more in #machine-learning 4 stories · sorted by recency
── more on @dust 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/dust-pretraining-tra…] indexed:0 read:28min 2026-10-05 · —