The paper introduces Dust, a zeroth‑order optimization algorithm that pretrains transformer language models without backpropagation. Dust adds independent Gaussian noise to the activation output of each linear layer at every token position, treats each token as a virtual population member, and estimates the gradient by averaging reward‑weighted noise over a population of draws, where the reward is the loss reduction attributable to that token’s perturbation. This activation‑space perturbation yields a population size proportional to the sequence length, allowing a single forward pass to evaluate thousands of virtual members in parallel. Experiments on GPT‑style models trained on FineWeb show that Dust’s test loss approaches that of backprop as the population grows; at 10 M–20 M tokens the gap narrows and Dust can surpass backprop with sufficiently large populations. Compared to weight‑space evolution strategies such as EGGROLL, Dust is 10³–10⁴ times more efficient from 1 M tokens upward, because it avoids materializing perturbed weights. Surprisingly, larger models are more population‑efficient: a 243 M‑parameter model outperforms a 120× smaller model across most population sizes, indicating that overparameterization enlarges the effective search space. Gradient estimates from Dust align increasingly with backprop’s gradients as population increases, with cosine similarity remaining stable up to 1 B tokens tested, suggesting scalability of the method.
Read original
hackernews