In plain words: They rewrite three popular training algorithms—AdaGrad, RMSProp, and Adam—as smooth equations describing how weights drift over time, keeping a running memory of past gradient sizes. Simulations and stability checks show these equations track the real step-by-step updates closely.
Abstract
In this paper, we propose a continuous-time formulation for the AdaGrad, RMSProp, and Adam optimization algorithms by modeling them as first-order integro-differential equations. We perform numerical simulations of these equations, along with stability and convergence analyses, to demonstrate their validity as accurate approximations of the original algorithms. Our results indicate a strong agreement between the behavior of the continuous-time models and the discrete implementations, thus providing a new perspective on the theoretical understanding of adaptive optimization methods.
Carlos Heredia
arXiv:2411.09734 · cs.LG, math.NA, math.OC · submitted Nov 14, 2024 · updated Jun 5, 2026
abstract · pdf · html · 60 pages, 15 figures; v3 - Section 4 corrected