Researchers have developed Dust, a method that pretrains transformer language models without relying on traditional backpropagation. By perturbing activations at each token, Dust efficiently estimates gradients in parallel, showing competitive performance and potential to surpass backprop in high-compute scenarios. The approach significantly outperforms existing evolutionary strategies in efficiency, especially at larger model scales.
qlabs.sh
· 2026-10-05
A technical blog post (originally from 2018) revisits a classic question about neural network training: why does the backpropagation algorithm invented by Rumelhart et al. in 1986 compute gradients by propagating errors backward through the network rather than forward, given that both directions are mathematically valid applications of the chain rule. The author walks through the notation and structure of a neural network node to set up an explanation of why forward-mode differentiation is computationally suboptimal compared to the standard backward approach.
gregorygundersen.com
· 2026-09-21
Sakana AI researchers Jeffrey Seely and Julian Gould introduced PC-ALM (Augmented Lagrangian Predictive Coding), a training method that replaces backpropagation's forward and backward passes with local feedback control dynamics running between neighboring layers. The team reports it successfully trains residual MLPs as deep as 1000 layers, reaching performance close to standard backpropagation while relying only on layer-local computation.
pub.sakana.ai
· 2026-09-14