Tech News
← Home  ·  All topics

Backpropagation

3 GoKawiil briefs on this topic

Dust introduces zeroth-order pretraining for transformers, rivaling backprop

Researchers have developed Dust, a method that pretrains transformer language models without relying on traditional backpropagation. By perturbing activations at each token, Dust efficiently estimates gradients in parallel, showing competitive performance and potential to surpass backprop in high-compute scenarios. The approach significantly outperforms existing evolutionary strategies in efficiency, especially at larger model scales.

Explainer revisits why backpropagation computes gradients backward, not forward

A technical blog post (originally from 2018) revisits a classic question about neural network training: why does the backpropagation algorithm invented by Rumelhart et al. in 1986 compute gradients by propagating errors backward through the network rather than forward, given that both directions are mathematically valid applications of the chain rule. The author walks through the notation and structure of a neural network node to set up an explanation of why forward-mode differentiation is computationally suboptimal compared to the standard backward approach.

Sakana AI unveils PC-ALM, a backprop-free method for training 1000-layer networks

Sakana AI researchers Jeffrey Seely and Julian Gould introduced PC-ALM (Augmented Lagrangian Predictive Coding), a training method that replaces backpropagation's forward and backward passes with local feedback control dynamics running between neighboring layers. The team reports it successfully trains residual MLPs as deep as 1000 layers, reaching performance close to standard backpropagation while relying only on layer-local computation.