In plain words: A textbook that explains neural network designs, training algorithms, and the math behind why they work, including how networks can solve equations from physics. It starts from zero background and aims to give students and practitioners a solid mathematical foundation.
Abstract
This book aims to provide an introduction to the topic of deep learning algorithms. We review essential components of deep learning algorithms in full mathematical detail including different artificial neural network (ANN) architectures (such as fully-connected feedforward ANNs, convolutional ANNs, recurrent ANNs, residual ANNs, and ANNs with batch normalization) and different optimization algorithms (such as the basic stochastic gradient descent (SGD) method, accelerated methods, and adaptive methods). We also cover several theoretical aspects of deep learning algorithms such as approximation capacities of ANNs (including a calculus for ANNs), optimization theory (including Kurdyka-Łojasiewicz inequalities), and generalization errors. In the last part of the book some deep learning approximation methods for PDEs are reviewed including physics-informed neural networks (PINNs) and deep Galerkin methods. We hope that this book will be useful for students and scientists who do not yet have any background in deep learning at all and would like to gain a solid foundation as well as for practitioners who would like to obtain a firmer mathematical understanding of the objects and methods considered in deep learning.
Arnulf Jentzen, Benno Kuckuck, Philippe von Wurstemberger
arXiv:2310.20360 · cs.LG, cs.AI, math.NA, math.PR, stat.ML · submitted Oct 31, 2023 · updated Jul 15, 2025
abstract · pdf · 737 pages, 33 figures, 45 source codes, 87 exercises. In v3, Chapters 5, 6, and 7 in Part III (Optimization) have been expanded
Look eg at the SGD chapter. I picked this because I think optimization is one of the areas where mathematicians actually can and do make impactful contributions to ML. But then look at the chapter in the book: most of the proofs are fairly elementary (like bias-variance decompositions or Jensen inequalities), some more interesting theorems (on convergence) are cited from the literature and do not build on the lemmata, and the sub-chapters on the actually interesting methods like ADAM,... are completely free of proofs or theory. It seems to me that after reading the chapter, a reader will have a good understanding of modern SGD methods and how we got there, but they won't necessarily be much wiser about why those methods work, other than having a good intuition confirmed by numerical experiments. If that's the outcome, then I wonder what the fuss proving all the basic stuff was all for. Wouldn't it be more useful to dedicate the space to convergence proofs for ADAM (which do exist) rather than showing lots of stuff like E(XY) = E(X)E(Y) for independent random variables?
That's just one chapter, I may not be doing them full justice here, although I did read through a few others as well. I first got this impression from the ANN chapter, which is ripe with long proofs for rather basic and uninteresting stuff, and from the physics-informed neural networks paper (which I actually find really nice, although it suffers a bit from the same problem as the SGD chapter). I don't want to be too critical here, it is nice in general to move towards a more rigorous and unified exposition of ML methods, and their approach should extend to the more technical results as well, just questioning where they drew the line of what to include and what not.