about
Layer rotation: a surprisingly powerful indicator of generalization? (arxiv.org)
2 points by MAXPOOL on Jul 2, 2019 | hide | past | pdf | discuss on HN

In plain words: Track how far each layer's weights drift from their starting values during training; that drift angle predicts how well the finished network handles new data. When every layer turns 90 degrees from its start, test accuracy rose by up to 30% versus other setups.

Abstract · Layer rotation: a surprisingly powerful indicator of generalization in deep networks?

Our work presents extensive empirical evidence that layer rotation, i.e. the evolution across training of the cosine distance between each layer's weight vector and its initialization, constitutes an impressively consistent indicator of generalization performance. In particular, larger cosine distances between final and initial weights of each layer consistently translate into better generalization performance of the final model. Interestingly, this relation admits a network independent optimum: training procedures during which all layers' weights reach a cosine distance of 1 from their initialization consistently outperform other configurations -by up to 30% test accuracy. Moreover, we show that layer rotations are easily monitored and controlled (helpful for hyperparameter tuning) and potentially provide a unified framework to explain the impact of learning rate tuning, weight decay, learning rate warmups and adaptive gradient methods on generalization and training speed. In an attempt to explain the surprising properties of layer rotation, we show on a 1-layer MLP trained on MNIST that layer rotation correlates with the degree to which features of intermediate layers have been trained.

Simon Carbonnelle, Christophe De Vleeschouwer
arXiv:1806.01603 · cs.LG, cs.CV, stat.ML · submitted Jun 5, 2018 · updated Jul 1, 2019
abstract · pdf · html · Extended version of paper presented at ICML workshop "Identifying and Understanding Deep Learning Phenomena"

add comment on HN
Also discussed: Jul 2019 (1 point, 0 comments)