about
Random feedback weights support learning in deep neural networks (arxiv.org)
77 points by plg on Nov 28, 2014 | hide | past | pdf | 14 comments on HN

In plain words: Instead of sending error signals back through the exact learned connections, this scheme sends them through random ones and lets the network learn to use them. It learned as quickly and accurately as the usual method that computes each neuron's share of the blame.

Abstract

The brain processes information through many layers of neurons. This deep architecture is representationally powerful, but it complicates learning by making it hard to identify the responsible neurons when a mistake is made. In machine learning, the backpropagation algorithm assigns blame to a neuron by computing exactly how it contributed to an error. To do this, it multiplies error signals by matrices consisting of all the synaptic weights on the neuron's axon and farther downstream. This operation requires a precisely choreographed transport of synaptic weight information, which is thought to be impossible in the brain. Here we present a surprisingly simple algorithm for deep learning, which assigns blame by multiplying error signals by random synaptic weights. We show that a network can learn to extract useful information from signals sent through these random feedback connections. In essence, the network learns to learn. We demonstrate that this new mechanism performs as quickly and accurately as backpropagation on a variety of problems and describe the principles which underlie its function. Our demonstration provides a plausible basis for how a neuron can be adapted using error signals generated at distal locations in the brain, and thus dispels long-held assumptions about the algorithmic constraints on learning in neural circuits.

Timothy P. Lillicrap, Daniel Cownden, Douglas B. Tweed, Colin J. Akerman
arXiv:1411.0247 · q-bio.NC, cs.NE · submitted Nov 2, 2014
abstract · pdf · html · 14 pages, 5 figures in main text; 13 pages appendix

add comment on HN

This reminds me of something that came up in Andrew Ng's online ML class. He said that it is important to check the correctness of your gradient calculation in backprop (by comparing it to a finite difference of the loss) because if you have a bug there, your algorithm might more or less work anyway, making it hard to tell that there was a bug. Apparently you can still get sensible output even with an incorrect gradient.
On reading this, my first question is about the properties of the "random" feedback matrix. They illustrate what is happening using a tiny 1-width machine and a "random" matrix of "1". It seems like some analysis needs to be done on what kind of "random" is most appropriate to replace the gradient update for larger machines. There could be something really interesting going on such that you could generate some optimal non-random B according to whatever the network topology is.
The implications of this are huge, it should drastically reduce processing time for neural nets. I wonder if given this if networks could be updated asynchronously/continuously.
I don't really understand how it would reduce processing time, could you elaborate?

The main implications seem to be for neuroscience, as far as I can tell. Backprop is considered biologically implausible because it requires either bidirectional communication over synapses (which doesn't happen) or weight sharing between neurons. But this allows the forward and backward connections to be decoupled (i.e. they are different synapses).

This is really interesting stuff, my first reaction was "why does this even work?" I think I still don't really fully understand what's going on.

> Backprop is considered biologically implausible

This is not true. See Neural Back propagation [1]. There are known mechanisms for backwards feedback between neural connections, for example Spike Timing Dependent Plasticity - where neural inputs that are well correlated in time and potential to output firings are strengthened over time. These phenomena are vital to learning and neural development.

[1] http://en.m.wikipedia.org/wiki/Neural_backpropagation

Yes but that's not really anything like the backpropagation algorithm in artificial neural networks.
From reading the abstract it seems that they are claiming that introducing some randomness into your gradient of weight changes allows for the quicker convergence of solution - I did not read the paper. I also don't exactly understand why it works - it sounds like they are claiming traditional back prop has room for improvement.
That's a very old strategy called jittering (also see stochastic gradient descent.)

This is something entirely different. They are not doing regular backpropagation at all, but somehow using neurons to learn how to backpropagate values. I haven't read the paper yet, just read their slides earlier, so that might not be correct.

Revolutionary paper indeed, if this idea generalizes to bigger deeper networks! (Then) why hasn't it been discovered before? It's like the Cambrian explosion of neuron networks, very exciting times.
okay but how does this save computing time? computing random matrices is not faster than (implicitly) transposing the weight matrix. so it only has philosophical implications, right?

Ok, got it: It will simplify the approach of how to create hardware based neural networks! no more complicated look-ups of the transposed weight matrix needed.

P = NP in the presence of a random oracle.
No it doesn't? What on Earth are you referring to?
keyword: oracle
random oracle - the random thing is what throws me...