about
Noise-Driven Escape from Metastable Phases Explains Grokking in DNNs (arxiv.org)
2 points by jerlendds 38 days ago | hide | past | pdf | 4 comments on HN

In plain words: A network can get stuck in a low-accuracy state, and only random jolts from training knock it over the barrier to a better one. This explains grokking: the model stays wrong for ages, then suddenly generalizes, with waits spanning two orders of magnitude.

Abstract · Noise-Driven Escape from Metastable Phases explains Grokking in Deep Neural Networks

Deep neural networks (DNNs) exhibit first order phase transitions under variations of the L2 regularization strength, with each transition marking the onset of a new learnable feature. Below a critical regularization strength, all features are in principle learnable, but coexisting metastable states, separated by energy barriers, can trap the network and impede convergence. A strength of DNNs is their ability to generalize. But many open questions remain, among them the origin of so called grokking: the abrupt, delayed onset of generalization after prolonged apparent overfitting. We show for linear DNNs that grokking is consistent with hysteresis in first-order L2 phase transitions: using L2 regularization to engineer deliberate trapping, we demonstrate that a model in a low-accuracy metastable state escapes only when SGD noise drives it across an energy barrier, with escape times following Arrhenius scaling. We reproduce grokking-like delayed convergence across two orders of magnitude in escape time by deliberately trapping models in metastable phases. Using sparse sub-sampling we also reproduce the canonical grokking curve where test error eventually approaches the final training error. Our work suggests that the number of metastable states equals the number of learnable features -- one per singular value of the data covariance -- the potential for hysteresis grows naturally with task complexity. We provide evidence that the same mechanism likely operates in general nonlinear DNNs. Our results provide routes toward more efficient learning schemes.

Ibrahim Talha Ersoy, Karoline Wiesner
arXiv:2606.17120 · cs.LG, physics.chem-ph · submitted Jun 15, 2026
abstract · pdf · html · 13 pages, 4 figures. Accepted at HiLD 2026: 4th Workshop on High-dimensional Learning Dynamics

add comment on HN

What’s a one sentence summary in language anyone can understand?
They're trying to explain the phenomenon, where during the training of a deep neural network, sometimes it will spend a long time with suboptimal parameters and then suddenly improve a lot very quickly, which has been termed "grokking." To do so, they study deep linear networks instead, which are easier to analyze. In that setting, it's possible to show that for more than two layers, there are two solutions for each feature, one where the network handles the feature correctly and one where it ignores it completely. However, that second solution has worse loss, so it is only metastable: if a path can be found to a region with better loss, the model will go there instead and "grok" the feature.

This path is provided by noisy parameter updates: the model learns from random samples of the training data, so every time it changes in a slightly different direction. Given enough time, those directions can randomly add up to a lucky escape path from the metastable solution. This is analogous to thermodynamics, where you can have a particle randomly bouncing around until it encounters another particle it can combine with in a lower-energy state, and the higher the temperature, the more quickly it happens. They empirically measure the relationship between escape time and temperature and find that its form agrees with the thermodynamic explanation, although their theoretical approximation of the energy barrier differs a lot from what the measurements imply, which they attribute to additional corrections necessary to make the approximation exact.

(In other words, it's impossible to summarize this in one sentence without first explaining a bunch of background information.)

I think you're saying that they explain “grokking” as noisy training updates eventually helping a deep network escape a suboptimal but temporarily stable solution and abruptly learn a feature, much like thermal fluctuations push a particle into a lower-energy state.

Does that capture the essence of what you said?

Yes.