about
How to train deep neural networks without backprop (arxiv.org)
1 point by ChrisCinelli on May 20, 2023 | hide | past | pdf | discuss on HN

In plain words: A brain-like alternative to backprop nudges the network and learns from the change, but noise grows with parameter count. Perturbing activations instead of weights and giving small parameter groups local losses fixes this, matching backprop on MNIST and CIFAR-10 and beating backprop-free methods on ImageNet.

Abstract · Scaling Forward Gradient With Local Losses

Forward gradient learning computes a noisy directional gradient and is a biologically plausible alternative to backprop for learning deep neural networks. However, the standard forward gradient algorithm, when applied naively, suffers from high variance when the number of parameters to be learned is large. In this paper, we propose a series of architectural and algorithmic modifications that together make forward gradient learning practical for standard deep learning benchmark tasks. We show that it is possible to substantially reduce the variance of the forward gradient estimator by applying perturbations to activations rather than weights. We further improve the scalability of forward gradient by introducing a large number of local greedy loss functions, each of which involves only a small number of learnable parameters, and a new MLPMixer-inspired architecture, LocalMixer, that is more suitable for local learning. Our approach matches backprop on MNIST and CIFAR-10 and significantly outperforms previously proposed backprop-free algorithms on ImageNet.

Mengye Ren, Simon Kornblith, Renjie Liao, Geoffrey Hinton
arXiv:2210.03310 · cs.LG, cs.CV, cs.NE · submitted Oct 7, 2022 · updated Mar 2, 2023
abstract · pdf · html · 31 pages, ICLR 2023

add comment on HN