about
Information Matrices and Generalization (arxiv.org)
2 points by sel1 on Jun 20, 2019 | hide | past | pdf | discuss on HN

In plain words: Training speed depends on how sharply the loss curves and how jumpy the gradients are; this study examines how the two interact instead of one at a time. It shows earlier work went astray by confusing three related matrices: curvature, gradient spread, and the Fisher matrix.

Abstract · On the interplay between noise and curvature and its effect on optimization and generalization

The speed at which one can minimize an expected loss using stochastic methods depends on two properties: the curvature of the loss and the variance of the gradients. While most previous works focus on one or the other of these properties, we explore how their interaction affects optimization speed. Further, as the ultimate goal is good generalization performance, we clarify how both curvature and noise are relevant to properly estimate the generalization gap. Realizing that the limitations of some existing works stems from a confusion between these matrices, we also clarify the distinction between the Fisher matrix, the Hessian, and the covariance matrix of the gradients.

Valentin Thomas, Fabian Pedregosa, Bart van Merriënboer, Pierre-Antoine Mangazol, Yoshua Bengio, Nicolas Le Roux
arXiv:1906.07774 · cs.LG, stat.ML · submitted Jun 18, 2019 · updated Apr 6, 2020
abstract · pdf · html · Accepted to AISTATS 2020

add comment on HN