about
Cautious Optimizers: Improving Training with One Line of Code (arxiv.org)
66 points by tosh on Mar 3, 2025 | hide | past | pdf | 2 comments on HN

In plain words: A one-line tweak to momentum-based optimizers like AdamW makes each weight move only when the update agrees with the gradient's direction, then rescales the step. This sped up large-language-model training and image classification consistently, with almost no extra tuning.

Abstract

AdamW has been the default optimizer for transformer pretraining. For many years, our community searched for faster and more stable optimizers with only constrained positive outcomes. In this work, we propose a \textbf{one-line modification in Pytorch} to any momentum-based optimizer, which we rename cautious optimizer, e.g. C-AdamW and C-Lion. Our theoretical result shows that this modification preserves Adam's Hamiltonian function and it does not break the convergence guarantee under the Lyapunov analysis. In addition, a whole new family of optimizers is revealed by our theoretical insight. Among them, we pick the simplest one for empirical experiments, showing not only consistent speed-up on LLM pretraining, but also image classification, with minimum extra tuning on hyperparameters. Code is available at https://github.com/kyleliang919/C-Optim.

Kaizhao Liang, Lizhang Chen, Bo Liu, Qiang Liu
arXiv:2411.16085 · cs.LG, cs.AI, cs.CL, cs.CV, cs.DM · submitted Nov 25, 2024 · updated Feb 15, 2026
abstract · pdf · html

add comment on HN
Also discussed: Nov 2024 (1 point, 1 comment)

Damn, this is a strikingly simple modification. Basically, modern deep learning optimizers typically calculate the update to the weights each step using some kind of momentum and/or LR scaling based on the running variance of the gradients. This means that, in theory, the actual "instantaneous" gradients from a particular backward pass might point in a different direction than the actual update the optimizer applies. The change the authors propose is to simply ignore any parameter updates proposed by the optimizer that have the opposite sign of the current gradient from the most recent backwards pass. They're essentially saying "only apply the long-term stabilized update where it agrees with the current 'instantaneous' gradient." They show that this simple change significantly speeds up model training.

I'm pretty intrigued by this, but will, as usual, wait for independent replications to come out before I fully believe it. That said, because of how simple this is, I'd expect such replications to happen within 24 hours. Exciting work!

I wonder if there mioght not be an opportunity for a warmup based mask inversion: for the first few epoches, only apply the momentum agreeing with instantaneous - after that, invert it since the momentum would technically have more info?

In any case, good idea - reminds me of the "apply same gradient multiple times" trick from a few years ago. May have weird behaviours at low batch sizes though...