In plain words: Replacing each weight with a tiny shared memory cell lets one network express learning algorithms, including backpropagation, just by running forward. It also finds new rules that work on data beyond its training range, learning by fast association instead of gradient descent.
Abstract · Meta Learning Backpropagation And Improving It
Many concepts have been proposed for meta learning with neural networks (NNs), e.g., NNs that learn to reprogram fast weights, Hebbian plasticity, learned learning rules, and meta recurrent NNs. Our Variable Shared Meta Learning (VSML) unifies the above and demonstrates that simple weight-sharing and sparsity in an NN is sufficient to express powerful learning algorithms (LAs) in a reusable fashion. A simple implementation of VSML where the weights of a neural network are replaced by tiny LSTMs allows for implementing the backpropagation LA solely by running in forward-mode. It can even meta learn new LAs that differ from online backpropagation and generalize to datasets outside of the meta training distribution without explicit gradient calculation. Introspection reveals that our meta learned LAs learn through fast association in a way that is qualitatively different from gradient descent.
Louis Kirsch, Jürgen Schmidhuber
arXiv:2012.14905 · cs.LG, cs.AI, cs.NE, stat.ML · submitted Dec 29, 2020 · updated Mar 13, 2022
abstract · pdf · html · Updated to the NeurIPS 2021 camera ready; fixed typo in eq 4
RNNs keep a memory of prior values such that you can pass on a “memory”.
At the end what this is doing is replacing components of the graph with mini-RNNs and pruning based on another network overseeing the first.
Having done quite a bit of work in this space I have a couple of thoughts.
What was the major advance in games? It was networks playing themselves.
Here we may want to do the same thing. Meta-learners need lots of “experience” (data) just like anything living.
This paper doesn’t dive too in-depth on the idea that the meta learner can be made general, but they can. There’s only so many problem types (classification, regression, generation, etc) and only so many data formats. Further, networks are fairly well defined, much more so than language.
This is the field of AutoML of anyones interested.