In plain words: Using the math of symmetries, each layer's pre-training step can be seen as a hunt for the simplest patterns—those that change the least under symmetry. This explains why networks learn simple features first and why features grow more complex in deeper layers.
Abstract · Why does Deep Learning work? - A perspective from Group Theory
Why does Deep Learning work? What representations does it capture? How do higher-order representations emerge? We study these questions from the perspective of group theory, thereby opening a new approach towards a theory of Deep learning. One factor behind the recent resurgence of the subject is a key algorithmic step called pre-training: first search for a good generative model for the input samples, and repeat the process one layer at a time. We show deeper implications of this simple principle, by establishing a connection with the interplay of orbits and stabilizers of group actions. Although the neural networks themselves may not form groups, we show the existence of {\em shadow} groups whose elements serve as close approximations. Over the shadow groups, the pre-training step, originally introduced as a mechanism to better initialize a network, becomes equivalent to a search for features with minimal orbits. Intuitively, these features are in a way the {\em simplest}. Which explains why a deep learning network learns simple features first. Next, we show how the same principle, when repeated in the deeper layers, can capture higher order representations, and why representation complexity increases as the layers get deeper.
Arnab Paul, Suresh Venkatasubramanian
arXiv:1412.6621 · cs.LG, cs.NE, stat.ML · submitted Dec 20, 2014 · updated Feb 28, 2015
abstract · pdf · html · 13 pages, 5 figures
+ An autoencoder is a stabilizer of the input f: it maps it to itself.
+ Imagine that the space of autoencoders forms a group.
+ Learning stops as soon as a stabilizer is found.
+ If the search is a Markov chain, the bigger the stabilizer the sooner it will be hit.
+ The group structure implies that this big stabilizer corresponds to a small orbit.
+ The orbit of an element x in X is the set of elements x can be moved to by the elements of G (the group).
+ The stabilizer is the set of group actions that map x to itself.
+ Blog post explaining fundamentals: https://gowers.wordpress.com/2011/11/09/group-actions-ii-the...
+ Reconstruction in deep NNs is often guided by an l2 distance. If there are competing feature sets, gradient descent moves the configuration to one of the stabilizers.
- Note. Is this indeed the case? Is there a WTA mechanism at play? Is this preferable?
- Note. In this view a NN is a MCMC that is stopped prematurely.
+ The probability that a network "discovers" a stabilizer for the signals f_i depends on the volume of the stabilizer.
- Would we be able to construct a PDF by running an MCMC till we are in a high probability region? Or does this only work if we're interested in modes?
+ A neural network operation is not a group action but can be seen as a "shadow group" in some auxiliary space (theorem 4.1, 4.2).
+ The size of the stabilizers is preserved in the original space (theorem 4.3).
+ Next levels exhibit again symmetry, generalized edges lead to figures like trapezoids, triangles, and butterflies.
Most important take-away: representations that exhibit a lot of symmetries are the ones that are easiest/fastest to find.