about
Why does unsupervised deep learning work? (arxiv.org)
90 points by MrQuincle on Jan 1, 2017 | hide | past | pdf | 9 comments on HN

In plain words: Using the math of symmetries, each layer's pre-training step can be seen as a hunt for the simplest patterns—those that change the least under symmetry. This explains why networks learn simple features first and why features grow more complex in deeper layers.

Abstract · Why does Deep Learning work? - A perspective from Group Theory

Why does Deep Learning work? What representations does it capture? How do higher-order representations emerge? We study these questions from the perspective of group theory, thereby opening a new approach towards a theory of Deep learning. One factor behind the recent resurgence of the subject is a key algorithmic step called pre-training: first search for a good generative model for the input samples, and repeat the process one layer at a time. We show deeper implications of this simple principle, by establishing a connection with the interplay of orbits and stabilizers of group actions. Although the neural networks themselves may not form groups, we show the existence of {\em shadow} groups whose elements serve as close approximations. Over the shadow groups, the pre-training step, originally introduced as a mechanism to better initialize a network, becomes equivalent to a search for features with minimal orbits. Intuitively, these features are in a way the {\em simplest}. Which explains why a deep learning network learns simple features first. Next, we show how the same principle, when repeated in the deeper layers, can capture higher order representations, and why representation complexity increases as the layers get deeper.

Arnab Paul, Suresh Venkatasubramanian
arXiv:1412.6621 · cs.LG, cs.NE, stat.ML · submitted Dec 20, 2014 · updated Feb 28, 2015
abstract · pdf · html · 13 pages, 5 figures

add comment on HN

To summarize:

+ An autoencoder is a stabilizer of the input f: it maps it to itself.

+ Imagine that the space of autoencoders forms a group.

+ Learning stops as soon as a stabilizer is found.

+ If the search is a Markov chain, the bigger the stabilizer the sooner it will be hit.

+ The group structure implies that this big stabilizer corresponds to a small orbit.

+ The orbit of an element x in X is the set of elements x can be moved to by the elements of G (the group).

+ The stabilizer is the set of group actions that map x to itself.

+ Blog post explaining fundamentals: https://gowers.wordpress.com/2011/11/09/group-actions-ii-the...

+ Reconstruction in deep NNs is often guided by an l2 distance. If there are competing feature sets, gradient descent moves the configuration to one of the stabilizers.

- Note. Is this indeed the case? Is there a WTA mechanism at play? Is this preferable?

- Note. In this view a NN is a MCMC that is stopped prematurely.

+ The probability that a network "discovers" a stabilizer for the signals f_i depends on the volume of the stabilizer.

- Would we be able to construct a PDF by running an MCMC till we are in a high probability region? Or does this only work if we're interested in modes?

+ A neural network operation is not a group action but can be seen as a "shadow group" in some auxiliary space (theorem 4.1, 4.2).

+ The size of the stabilizers is preserved in the original space (theorem 4.3).

+ Next levels exhibit again symmetry, generalized edges lead to figures like trapezoids, triangles, and butterflies.

Most important take-away: representations that exhibit a lot of symmetries are the ones that are easiest/fastest to find.

This could be the least dismissive and anti-intellectual tl;dr I've seen here.
I can't possibly be the only one here who wishes articles were presented in this format, with each of the points clickable for more details, such that the whole thing is a tree structure of summarized points.
> + Reconstruction in deep NNs is often guided by an l2 distance. If there are competing feature sets, gradient descent moves the configuration to one of the stabilizers. > - Note. Is this indeed the case? Is there a WTA mechanism at play? Is this preferable? > - Note. In this view a NN is a MCMC that is stopped prematurely.

Generally the 'performance' of the autoencoder is measured at the reconstruction step, i.e. what is the reconstruction accuracy of the autoencoder? This is an L2 loss function.

This is pretty cool; the link between machine learning and group theory is neat. Do you know if this link is new? I've not heard of it before. Thanks for the summary!
This particular link seems new indeed.

There have been other links listed on HN in the past. One that came back a few times is on renormalization group theory (actually a semigroup).

Currently my new angle of attack is to check what people have to say who come from a representation theory angle.

I also liked this paper by Persi Diaconis a lot: http://statweb.stanford.edu/~cgates/PERSI/papers/The%20MCMC%...

I also like that on his wikipedia page (https://www.wikiwand.com/en/Persi_Diaconis) it's stated that: According to Martin Gardner, at school, Diaconis supported himself by playing poker on ships between New York and South America. Gardner recalls that Diaconis had "fantastic second deal and bottom deal". :-)

You are my hero.
Will deep learning soon make the jump from alchemy to chemistry? Will abstract algebra suddenly attract all the business-brogrammers? Stay tuned!
To be clear, this paper looks great. My concern is the poeple pushing deep learning before the theory is worked out—it's an AI winter waiting to happen.