about
Predictive Coding Approximates Backprop Along Arbitrary Computation Graphs (2020) (arxiv.org)
65 points by DanielBMarkham on May 2, 2021 | hide | past | pdf | 9 comments on HN

In plain words: A brain-like learning rule, where each unit only adjusts itself and its neighbors, now works on any network of connected steps, not just simple layered ones. It settles on the same error signals as the standard whole-network training algorithm and matches its accuracy.

Abstract · Predictive Coding Approximates Backprop along Arbitrary Computation Graphs

Backpropagation of error (backprop) is a powerful algorithm for training machine learning architectures through end-to-end differentiation. However, backprop is often criticised for lacking biological plausibility. Recently, it has been shown that backprop in multilayer-perceptrons (MLPs) can be approximated using predictive coding, a biologically-plausible process theory of cortical computation which relies only on local and Hebbian updates. The power of backprop, however, lies not in its instantiation in MLPs, but rather in the concept of automatic differentiation which allows for the optimisation of any differentiable program expressed as a computation graph. Here, we demonstrate that predictive coding converges asymptotically (and in practice rapidly) to exact backprop gradients on arbitrary computation graphs using only local learning rules. We apply this result to develop a straightforward strategy to translate core machine learning architectures into their predictive coding equivalents. We construct predictive coding CNNs, RNNs, and the more complex LSTMs, which include a non-layer-like branching internal graph structure and multiplicative interactions. Our models perform equivalently to backprop on challenging machine learning benchmarks, while utilising only local and (mostly) Hebbian plasticity. Our method raises the potential that standard machine learning algorithms could in principle be directly implemented in neural circuitry, and may also contribute to the development of completely distributed neuromorphic architectures.

Beren Millidge, Alexander Tschantz, Christopher L. Buckley
arXiv:2006.04182 · cs.LG, cs.NE · submitted Jun 7, 2020 · updated Oct 5, 2020
abstract · pdf · html · Submitted to NeurIPS 2020. Updated Acknowledgements. 11/06/20: fixed typos in maths -- 11/07/20: minor corrections; 05/10/20: major rewrite for ICLR

add comment on HN

Also, most recent state of progress: Predictive Coding Can Do Exact Backpropagation on Any Neural Network (2021)

https://arxiv.org/abs/2103.04689

here's a link with reviewer comments https://openreview.net/forum?id=PdauS7wZBfC (praise to openreview!)
The decision reasoning is super helpful to put things in context. The arbitrary binary decision to "accept" vs "reject" especially for the snooty "high bar for acceptance at ICLR" is laughable in a world of free information access.
AstralCodexTen (formerly SlateStarCodex) has discussed this here - https://astralcodexten.substack.com/p/link-unifying-predicti...

He mostly points to this post in LessWrong - https://www.lesswrong.com/posts/JZZENevaLzLLeC3zn/predictive...

If backprop is not needed, would this finding make automatic-differentiation functionality obsolete in DL frameworks, allowing these frameworks to become much simpler? Or is there still some constant factor that makes backprop favorable?
Reading the openreview link, the current understanding is that this approach is dramatically more computationally intensive than standard backprop - limiting its utility.
Backprop is simpler.
Needs [2020] in the title. Interesting work nevertheless.
This (and its follow up papers) have already been discussed multiple times here. Don't have the links handy though..