In plain words: A proof shows gradient training reaches zero training loss in a wide deep network with skip connections, in steps growing only moderately with size. A matrix of how training examples interact stays stable as weights update; the same holds for deep residual convolutional networks.
Abstract
Gradient descent finds a global minimum in training deep neural networks despite the objective function being non-convex. The current paper proves gradient descent achieves zero training loss in polynomial time for a deep over-parameterized neural network with residual connections (ResNet). Our analysis relies on the particular structure of the Gram matrix induced by the neural network architecture. This structure allows us to show the Gram matrix is stable throughout the training process and this stability implies the global optimality of the gradient descent algorithm. We further extend our analysis to deep residual convolutional neural networks and obtain a similar convergence result.
Simon S. Du, Jason D. Lee, Haochuan Li, Liwei Wang, Xiyu Zhai
arXiv:1811.03804 · cs.LG, cs.AI, cs.CV, math.OC, stat.ML · submitted Nov 9, 2018 · updated May 28, 2019
abstract · pdf · html · ICML 2019
The major contribution of the work is showing that ResNet needs a number of parameters which is polynomial in the dataset size to converge to a global optimum in contrast to traditional neural nets which require an exponential number of parameters.