In plain words: Adam, the standard rule that nudges weights during training, can slip into a state where its updates grow large and point randomly instead of downhill, so the loss blows up. Runs from 7 to 546 billion parameters show big batches make this most likely.
Abstract
We present a theory for the previously unexplained divergent behavior noticed in the training of large language models. We argue that the phenomenon is an artifact of the dominant optimization algorithm used for training, called Adam. We observe that Adam can enter a state in which the parameter update vector has a relatively large norm and is essentially uncorrelated with the direction of descent on the training loss landscape, leading to divergence. This artifact is more likely to be observed in the training of a deep model with a large batch size, which is the typical setting of large-scale language model training. To argue the theory, we present observations from the training runs of the language models of different scales: 7 billion, 30 billion, 65 billion, and 546 billion parameters.
Igor Molybog, Peter Albert, Moya Chen, Zachary DeVito, David Esiobu, Naman Goyal, Punit Singh Koura, Sharan Narang, Andrew Poulton, Ruan Silva, Binh Tang, Diana Liskovich, et al.
arXiv:2304.09871 · cs.LG, cs.AI, math.OC · submitted Apr 19, 2023 · updated Apr 25, 2023
abstract · pdf · html