Tech News
← Home  ·  All topics

Eggroll

1 GoKawiil brief on this topic

Dust introduces zeroth-order pretraining for transformers, rivaling backprop

Researchers have developed Dust, a method that pretrains transformer language models without relying on traditional backpropagation. By perturbing activations at each token, Dust efficiently estimates gradients in parallel, showing competitive performance and potential to surpass backprop in high-compute scenarios. The approach significantly outperforms existing evolutionary strategies in efficiency, especially at larger model scales.