Dust Explores Transformer Pretraining Without Backpropagation
Introduction
Backpropagation has shaped nearly every modern deep-learning system. It exploits differentiability to assign credit through a network and gives optimizers a direct gradient signal. Dust asks whether that assumption could eventually be relaxed: with enough compute, might broad search-based credit assignment become competitive with backpropagation?
How Dust works
Dust is a zeroth-order optimization method. Rather than analytically differentiating the loss, it perturbs internal activations, measures how those perturbations affect the loss, and aggregates the perturbations according to their rewards. Perturbations that reduce the loss receive a more useful contribution to the estimated update direction.
The central design choice is to operate in activation space rather than weight space. Dust applies perturbations independently at every token, treating tokens as members of a virtual population. A single forward pass can therefore evaluate many candidates in parallel, without materializing a separate copy of the model for each candidate. The approach also uses different token-level credit rules for different transformer layer types and includes measures intended to limit interference between perturbed modules.
Key findings
- The authors describe Dust as one of the first zeroth-order approaches to approach backpropagation on transformer language-model pretraining.
- As the virtual population grows, its gradient estimates become more aligned with backpropagation. In several tested settings, Dust reportedly surpasses it.
- Activation-space search avoids much of the cost associated with weight-space evolutionary strategies. The paper estimates that, beyond a certain token scale, Dust can be several orders of magnitude more efficient than the transformer implementation of EGGROLL used as a comparison.
- Larger models were often more population-efficient rather than less. In the reported experiments, a substantially larger model outperformed a much smaller one across many population sizes.
- This is not yet a replacement for backpropagation. The authors explicitly frame the work as a foundation and do not claim that Dust is currently compute-efficient enough for routine use.
Why it matters
The significance of Dust is less about an immediate training recipe and more about expanding the set of plausible learning algorithms for large neural networks. The results suggest that zeroth-order methods need not be restricted to small systems or simple objectives. With parallel activation perturbations, they can be tested against one of the hardest current workloads: transformer pretraining.
Activation-space search may also make it easier to train systems containing non-differentiable components, external programs, or other structures that do not fit neatly into conventional computational graphs. It offers a way to search over internal representations rather than only over parameters.
The limitations are equally important. Matching backpropagation can require a larger population and substantially more computation. The reported evidence demonstrates feasibility and encouraging scaling behavior, not a general win in cost, data efficiency, or final capability. Future work will need to test whether the approach remains competitive at larger scales and in architectures that genuinely benefit from non-differentiable components.
Source: Hacker News
Comments
Checking sign-in status...
Loading comments...