about
Sophia: Scalable Stochastic 2nd-Order Optimizer for Language Model Pre-Training (arxiv.org)
54 points by tosh on Apr 7, 2024 | hide | past | pdf | 2 comments on HN

In plain words: It divides each parameter's recent gradient by a rough estimate of how sharply the loss curves there, refreshed only every few steps and capped in size to stay stable. Training GPT models, it matched Adam's quality in half the steps, halving compute and wall-clock time.

Abstract · Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Given the massive cost of language model pre-training, a non-trivial improvement of the optimization algorithm would lead to a material reduction on the time and cost of training. Adam and its variants have been state-of-the-art for years, and more sophisticated second-order (Hessian-based) optimizers often incur too much per-step overhead. In this paper, we propose Sophia, Second-order Clipped Stochastic Optimization, a simple scalable second-order optimizer that uses a light-weight estimate of the diagonal Hessian as the pre-conditioner. The update is the moving average of the gradients divided by the moving average of the estimated Hessian, followed by element-wise clipping. The clipping controls the worst-case update size and tames the negative impact of non-convexity and rapid change of Hessian along the trajectory. Sophia only estimates the diagonal Hessian every handful of iterations, which has negligible average per-step time and memory overhead. On language modeling with GPT models of sizes ranging from 125M to 1.5B, Sophia achieves a 2x speed-up compared to Adam in the number of steps, total compute, and wall-clock time, achieving the same perplexity with 50% fewer steps, less total compute, and reduced wall-clock time. Theoretically, we show that Sophia, in a much simplified setting, adapts to the heterogeneous curvatures in different parameter dimensions, and thus has a run-time bound that does not depend on the condition number of the loss.

Hong Liu, Zhiyuan Li, David Hall, Percy Liang, Tengyu Ma
arXiv:2305.14342 · cs.LG, cs.CL, math.OC · submitted May 23, 2023 · updated Mar 5, 2024
abstract · pdf · html

add comment on HN
Also discussed: Apr 2026 (4 points, 1 comment) · Jul 2023 (5 points, 0 comments) · May 2023 (7 points, 0 comments) · May 2023 (2 points, 0 comments)

I wonder how it relates/compares to https://arxiv.org/abs/2205.08253 : Title: Momentum-Based Policy Gradient with Second-Order Information

Abstract: Variance-reduced gradient estimators for policy gradient methods have been one of the main focus of research in the reinforcement learning in recent years as they allow acceleration of the estimation process. We propose a variance-reduced policy-gradient method, called SHARP, which incorporates second-order information into stochastic gradient descent (SGD) using momentum with a time-varying learning rate. SHARP algorithm is parameter-free, achieving ϵ-approximate first-order stationary point with O(ϵ−3) number of trajectories, while using a batch size of O(1) at each iteration. Unlike most previous work, our proposed algorithm does not require importance sampling which can compromise the advantage of variance reduction process. Moreover, the variance of estimation error decays with the fast rate of O(1/t2/3) where t is the number of iterations. Our extensive experimental evaluations show the effectiveness of the proposed algorithm on various control tasks and its advantage over the state of the art in practice.

Though I guess it may be more suitable to training on more interactive tasks like those emphasizing in-context learning, to better exploit RL's ability to adequately deal with tasks where early parts of the model's output are left open, especially when used with curriculum learning to gradually build up capability. Might be more suited to diffusion models than classic GPT 's, though, as RL shines where models have more agency, even if it's just deciding in what order to write the output.

I'm not too familiar with SHARP but I've implemented SophiaH before, it is effectively mSGD but with the momentum being scaled by inverse of a separate hessian moment with some additional clipping. It works surprisingly well if you can afford the additional memory/compute.

It seems like SHARP directly incorporates the hessian directly into the SGD momentum, and increases the 'alpha' or momentum contribution over time.