about
Unreasonable Effectivness of Deep Learning (arxiv.org)
1 point by ghosthamlet on Apr 5, 2018 | hide | past | pdf | discuss on HN

In plain words: Averaging small memory machines by weighted voting turns out to do the same math as backpropagation, the usual way to train a neural network. This match lets known guarantees from online statistics bound how well deep learning predicts, and what different network shapes can compute.

Abstract

We show how well known rules of back propagation arise from a weighted combination of finite automata. By redefining a finite automata as a predictor we combine the set of all $k$-state finite automata using a weighted majority algorithm. This aggregated prediction algorithm can be simplified using symmetry, and we prove the equivalence of an algorithm that does this. We demonstrate that this algorithm is equivalent to a form of a back propagation acting in a completely connected $k$-node neural network. Thus the use of the weighted majority algorithm allows a bound on the general performance of deep learning approaches to prediction via known results from online statistics. The presented framework opens more detailed questions about network topology; it is a bridge to the well studied techniques of semigroup theory and applying these techniques to answer what specific network topologies are capable of predicting. This informs both the design of artificial networks and the exploration of neuroscience models.

Finn Macleod
arXiv:1803.10768 · cs.LG, stat.ML · submitted Mar 28, 2018
abstract · pdf · html

add comment on HN
Also discussed: Apr 2018 (3 points, 0 comments)