In plain words: A survey collects new mathematical tools that explain deep learning puzzles classical learning theory cannot, such as why huge networks generalize, why training finds good solutions despite many traps, and what depth adds. The tools give partial answers to these questions.
Abstract
We describe the new field of mathematical analysis of deep learning. This field emerged around a list of research questions that were not answered within the classical framework of learning theory. These questions concern: the outstanding generalization power of overparametrized neural networks, the role of depth in deep architectures, the apparent absence of the curse of dimensionality, the surprisingly successful optimization performance despite the non-convexity of the problem, understanding what features are learned, why deep architectures perform exceptionally well in physical problems, and which fine aspects of an architecture affect the behavior of a learning task in which way. We present an overview of modern approaches that yield partial answers to these questions. For selected approaches, we describe the main ideas in more detail.
Julius Berner, Philipp Grohs, Gitta Kutyniok, Philipp Petersen
arXiv:2105.04026 · cs.LG, stat.ML · submitted May 9, 2021 · updated Feb 8, 2023
abstract · pdf · html · A version of this review paper appears as a chapter in the book "Mathematical Aspects of Deep Learning" by Cambridge University Press
IE, Deep learning is fundamentally just about getting the mathematically simple but complex and multi-layerd "neural networks" to do stuff. Training them, testing them and deploying them. There are many intuitions about these things but there's no complete theory - some intuitions involve mathematical analogies and simplifications while other involve "folk knowledge" or large scale experiments. And that's not saying folks giving math about deep learning aren't proving real things. It's just they characterizing the whole or even a substantial part of such systems.
It's not surprising that a complex like a many-layered Relu network can't fully characterized or solved mathematically. You'd expect that of any arbitrarily complex algorithmic construct. Differential equations of many variables and arbitrary functions also can't have their solutions fully characterized.