In plain words: Gradient-based training, the standard approach, can fail when the system is chaotic, since tiny changes grow and the gradient signal becomes useless. The report traces this to the system's Jacobian — the matrix of how outputs respond to inputs — and gives rules for spotting it.
Abstract · Gradients are Not All You Need
Differentiable programming techniques are widely used in the community and are responsible for the machine learning renaissance of the past several decades. While these methods are powerful, they have limits. In this short report, we discuss a common chaos based failure mode which appears in a variety of differentiable circumstances, ranging from recurrent neural networks and numerical physics simulation to training learned optimizers. We trace this failure to the spectrum of the Jacobian of the system under study, and provide criteria for when a practitioner might expect this failure to spoil their differentiation based optimization algorithms.
Luke Metz, C. Daniel Freeman, Samuel S. Schoenholz, Tal Kachman
arXiv:2111.05803 · cs.LG, stat.ML · submitted Nov 10, 2021 · updated Jan 21, 2022
abstract · pdf · html