cd /news/machine-learning/backprop-has-a-real-challenger-dust-… · home › topics › machine-learning › article
[ARTICLE · art-145888] src=byteiota.com ↗ pub= topic=machine-learning verified=true sentiment=· neutral

Backprop Has a Real Challenger: Dust Trains Transformers

Researchers at QLabs published Dust, a zeroth-order optimization method that matches backpropagation on transformer pretraining at scale, reaching a test loss of approximately 5.17 at 1 million tokens with 16,000 population draws versus backprop's 5.18. According to the Dust paper, the method adds independent Gaussian noise to activation outputs in a single forward pass, making it 1,000 to 10,000 times more compute-efficient than EGGROLL, the previous state-of-the-art weight-space evolution strategy, though the authors state Dust needs orders of magnitude more compute efficiency before it is a practical replacement. Testing models from 2 million to 243 million parameters at a fixed 10 million tokens, the 243M model was more population-efficient than the 2M model at every population size above the smallest tested, which lead author Samip Dahal cited in saying "many of the core assumptions in optimization research are completely wrong.

read4 min views1 publishedOct 6, 2026
Backprop Has a Real Challenger: Dust Trains Transformers
Image: Byteiota (auto-discovered)

Researchers at QLabs published Dust this week — the first zeroth-order optimization method competitive with backpropagation at pretraining transformer language models. The paper, trending on Hacker News today with 153+ points, makes a claim that challenges 40 years of deep learning orthodoxy: you do not need the chain rule to train a transformer. Backpropagation has never had a serious challenger. Dust is the first one that actually shows up with numbers.

To be clear: Dust’s authors are not claiming backprop is dead. They explicitly write that Dust needs “orders of magnitude more compute efficiency” before it’s a practical replacement. But “competitive at training scale” is a milestone nobody thought zeroth-order optimization could reach. That changes the conversation.

The Virtual Population Trick #

Standard evolution strategies train models by generating multiple copies of the network, perturbing each copy’s weights, measuring which perturbations lower the loss, and averaging the signal into a gradient estimate. This requires materializing full network copies — extremely expensive at scale.

Dust sidesteps this entirely with a “virtual population” concept. Instead of perturbing weights across separate copies, it adds independent Gaussian noise to activation outputs at every token position in a single forward pass. A 4,096-token sequence becomes a population of 4,096 virtual members, all evaluated in parallel — no extra copies needed. The loss reduction per token provides the reward signal. Weight gradients are then estimated from reward-weighted perturbations, producing an outer product form identical to backpropagation’s — except no chain rule was involved. According to the Dust paper, this makes it 1,000 to 10,000 times more compute-efficient than EGGROLL, the previous state-of-the-art weight-space evolution strategy.

What the Benchmarks Actually Show #

At small compute budgets, Dust trails backprop. At 100k tokens, performance is comparable but not clearly better. The story changes at scale. At 1 million tokens, Dust with 16,000 population draws achieves a test loss of approximately 5.17 — backprop lands at 5.18. Essentially tied. At 10 million tokens with the Adam optimizer, Dust’s power-law extrapolation projects a limit of 5.248 against backprop’s 5.361. Dust marginally ahead.

The honest caveat: 16,000 draws means 16,000 perturbation samples per weight update. That is computationally expensive by current standards. Researchers in the community have raised a legitimate methodological point — comparing by tokens rather than GPU-hours may flatter Dust. At equal GPU-hours, the gap may look different. The paper does not claim compute efficiency today; it claims it has cleared the conceptual barrier. Those are different things, and conflating them is how overhyped AI papers happen.

Related: arXiv Caps Submissions: AI Flood Breaks Science Publishing

The Counterintuitive Scaling Result #

The most surprising finding in Dust is not the benchmark performance — it is the scaling behavior. Standard intuition says zeroth-order methods get harder as models grow: more parameters means noisier gradient estimates. Dust shows the opposite.

Testing models from 2 million to 243 million parameters at fixed 10 million tokens, the 243M model is more population-efficient than the 2M model at every population size above the smallest tested. The authors interpret this as larger models having better-conditioned loss landscapes — more favorable geometry for search-based optimization. If this holds at GPT-4 scale, the efficiency argument in the paper gets stronger as models get bigger, not weaker. As lead author Samip Dahal noted on X, “many of the core assumptions in optimization research are completely wrong.” That is an extraordinary property, and the mechanism is not yet understood.

Why This Matters Beyond Backpropagation Benchmarks #

Memory savings are interesting. The architectural implications are more significant. Backpropagation requires end-to-end differentiability — everything in the training graph must have a well-defined gradient. This is a hard constraint. Transformers that interact with external programs, databases, or code interpreters cannot be trained with backprop directly; they require workarounds like REINFORCE or PPO that add complexity and instability.

Dust has no such requirement. The paper’s future work section explicitly names “nets with an external program in the loop or transformers looped over many steps” as architectures backpropagation struggles with but Dust could handle. Coverage from Digg noted the research community sees this as unlocking training for hardware that only runs forward passes — neuromorphic and analog chips where backprop is physically impractical. None of this is available today. However, the research establishes that the ceiling on these approaches is not theoretical — it is practical and solvable.

Key Takeaways #

  • Dust is the first zeroth-order optimizer to match backpropagation performance at transformer pretraining scale — a 40-year barrier cleared, at research scale
  • The virtual population concept makes it feasible: one forward pass evaluates thousands of population members in parallel, no network copies required
  • Larger models are more population-efficient in Dust, not less — the opposite of prior assumptions, and the mechanism is still unexplained
  • The authors are explicit: Dust is not ready to replace backprop today, and significant compute efficiency work remains before it could
  • The real long-term value is architectural freedom — training models containing non-differentiable components that backprop fundamentally cannot handle
── more in #machine-learning 4 stories · sorted by recency
── more on @qlabs 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/backprop-has-a-real-…] indexed:0 read:4min 2026-10-06 · —