about
Representation Benefits of Deep Feedforward Networks (arxiv.org)
1 point by chriskanan on Oct 20, 2015 | hide | past | pdf | discuss on HN

In plain words: A family of classification tasks is built where a network with just one layer of nodes needs exponentially many of them and still can't get error below 1/6. A deep network with only two nodes in each of 2k layers gets every case right.

Abstract

This note provides a family of classification problems, indexed by a positive integer $k$, where all shallow networks with fewer than exponentially (in $k$) many nodes exhibit error at least $1/6$, whereas a deep network with 2 nodes in each of $2k$ layers achieves zero error, as does a recurrent network with 3 distinct nodes iterated $k$ times. The proof is elementary, and the networks are standard feedforward networks with ReLU (Rectified Linear Unit) nonlinearities.

Matus Telgarsky
arXiv:1509.08101 · cs.LG, cs.NE · submitted Sep 27, 2015 · updated Sep 29, 2015
abstract · pdf · html

add comment on HN