The 40-Year Rule AI Researchers Want to Break
Every AI model, from large language models to image classifiers, trains using backpropagation, an algorithm popularized in 1986. This method works by running data forward through a neural network to make predictions, calculating a "loss" to measure error, and then using the calculus chain rule to determine precise gradients—how each weight should adjust to reduce that error.
Backpropagation's power lies in its efficient computation of these exact gradients, but it demands that every operation within the model be differentiable. This constraint often necessitates storing intermediate activations for the backward pass, influencing nearly every architectural choice and hardware design in deep learning today. The alternative, evolutionary strategies, involves guessing weight adjustments and iteratively checking for loss reduction; however, these methods are notoriously slow due to the need for multiple forward passes.
Researchers at Q Labs are challenging this 40-year paradigm with Dust, a novel training approach for transformers. Dust investigates whether a transformer can learn effectively from trial and error, not gradient calculation, by perturbing activations rather than weights. This research probes the fundamental question of whether machine learning can break free from the differentiability requirement, potentially enabling entirely new, non-differentiable model architectures.
Dust Turns Every Token Into an Experiment
Evolution strategies have long offered an alternative to backpropagation, but they traditionally operate in weight space. This means they perturb a model’s parameters directly, requiring a separate forward pass for each set of perturbed weights to evaluate its impact on the loss. Such methods are "painfully slow" due to the computational overhead of these individual evaluations.
Dust, a novel method from Q Labs, flips this paradigm by injecting independent random noise into a model’s activation space—specifically, into the output of each layer at every token position. Instead of guessing at weight changes, Dust estimates which activation perturbations reduce the loss for each token, leveraging these insights to infer appropriate weight adjustments.
This approach creates a "virtual population." A single forward pass with a 2,048-token sequence effectively runs 2,048 parallel perturbation trials. The noise that lowers the loss at a given token position is "rewarded," and these reward-weighted signals are then averaged to estimate the gradient for each layer.
From these estimated gradients, weight updates are computed using the same mechanism as backpropagation. The critical difference is that Dust derives its gradient estimates from trial and error within the activation space, rather than from calculus via the chain rule. This innovative method represents a significant departure from conventional training paradigms.
The Results Were Weird—and Worth Watching
Initial results from Q Labs revealed a bounded win for Dust. On shorter training runs (100k and 1M tokens), Dust surpassed tuned backpropagation in achieving a smarter model. While it didn't fully catch up on longer runs, more samples progressively narrowed the performance gap, with an extrapolated power-law fitted loss of 4.43 for Dust versus 4.63 for backpropagation at 20M tokens under large population budgets.
Comparing Dust to existing zeroth-order methods, the distinction was substantial. Dust performed significantly better than EGGROLL, even when EGGROLL received 256 times more population samples. This demonstrated Dust's superior efficiency in estimating gradients through activation perturbations.
Perhaps most counterintuitively, larger tested models appeared more population-efficient. This scaling benefit extended to models up to 243 million parameters. These experiments, however, used small GPT-style models based on modded-nanoGPT and the FineWeb dataset, not the production-scale LLMs we typically encounter. For further technical details, explore Dust: Pretraining Transformers Without Backpropagation - Q Labs.
Enjoying this? Get one like it in your inbox each morning.
one email a day · unsubscribe in two clicks · no third-party tracking
The Catch: A Clever Method That Costs Too Much
Dust’s ingenuity comes with a considerable computational cost. The token-level parallelism, while innovative, does not reduce total GPU-hours or FLOPs compared to backpropagation; it merely reallocates them. Q Labs’ researchers explicitly state that Dust requires orders of magnitude more compute, making it impractical as a direct replacement for backpropagation today, especially for models exceeding the 243 million parameters tested.
Near-term applications for derivative-free training methods like Dust are likely in areas where backpropagation struggles due to its differentiability requirement. This includes training models with non-differentiable components such as:
- External programs
- Complex simulators
- Discrete decision-making modules
- Repeated self-calls within the model architecture
Watch for efficiency improvements and the development of hybrid training methods that combine Dust’s flexibility with backpropagation’s speed. Such approaches could unlock novel architectures currently constrained by analytical differentiability. For those eager to inspect the mechanics, Q Labs provides detailed research on their website and an open-source implementation on GitHub.
Frequently Asked Questions
What is Dust?
Dust is an experimental method for training transformer models by perturbing activations and estimating learning directions from changes in loss, without using backpropagation.
How does Dust avoid backpropagation?
It adds random noise to layer activations across token positions, measures how the loss responds, and combines those signals to estimate updates.
Does Dust outperform backpropagation?
It beat tuned backpropagation in some short training runs, but it has not shown an overall compute-efficiency advantage at practical training scales.
Why could training without backpropagation matter?
It could make it easier to train systems involving non-differentiable operations, such as external programs or simulators, that do not fit neatly into standard gradient-based training.

