about
Rotational Equilibrium: How Weight Decay Balances Learning Across NeuralNetworks (arxiv.org)
2 points by veryluckyxyz on Mar 25, 2024 | hide | past | pdf | discuss on HN

In plain words: Weight decay settles each neuron's weight vector into a steady spin speed, making effective learning rates similar across layers and neurons. Setting that spin directly delivers weight decay's benefits while sharply cutting the need for learning rate warmup.

Abstract · Rotational Equilibrium: How Weight Decay Balances Learning Across Neural Networks

This study investigates how weight decay affects the update behavior of individual neurons in deep neural networks through a combination of applied analysis and experimentation. Weight decay can cause the expected magnitude and angular updates of a neuron's weight vector to converge to a steady state we call rotational equilibrium. These states can be highly homogeneous, effectively balancing the average rotation -- a proxy for the effective learning rate -- across different layers and neurons. Our work analyzes these dynamics across optimizers like Adam, Lion, and SGD with momentum, offering a new simple perspective on training that elucidates the efficacy of widely used but poorly understood methods in deep learning. We demonstrate how balanced rotation plays a key role in the effectiveness of normalization like Weight Standardization, as well as that of AdamW over Adam with L2-regularization. Finally, we show that explicitly controlling the rotation provides the benefits of weight decay while substantially reducing the need for learning rate warmup.

Atli Kosson, Bettina Messmer, Martin Jaggi
arXiv:2305.17212 · cs.LG · submitted May 26, 2023 · updated Jun 3, 2024
abstract · pdf · html · Accepted to ICML 2024; Code available at https://github.com/epfml/REQ

add comment on HN