In plain words: They count the neurons needed to represent polynomials with many variables as the variable count grows. A one-hidden-layer network needs exponentially more as variables are added, while deeper ones need only a linear increase, with each extra layer shrinking the exponent.
Abstract
It is well-known that neural networks are universal approximators, but that deeper networks tend in practice to be more powerful than shallower ones. We shed light on this by proving that the total number of neurons $m$ required to approximate natural classes of multivariate polynomials of $n$ variables grows only linearly with $n$ for deep neural networks, but grows exponentially when merely a single hidden layer is allowed. We also provide evidence that when the number of hidden layers is increased from $1$ to $k$, the neuron requirement grows exponentially not with $n$ but with $n^{1/k}$, suggesting that the minimum number of layers required for practical expressibility grows only logarithmically with $n$.
David Rolnick, Max Tegmark
arXiv:1705.05502 · cs.LG, cs.NE, stat.ML · submitted May 16, 2017 · updated Apr 27, 2018
abstract · pdf · html · Replaced to match version published at ICLR 2018. 14 pages, 2 figs