about
SoftServe: A Scalable Quasi-Newton Method for Deep Learning (arxiv.org)
1 point by E-Reverance 21 hours ago | hide | past | pdf | discuss on HN

In plain words: A new optimizer estimates how curved the loss landscape is, keeping those estimates positive even when the surface bends the wrong way, and uses fast matrix multiplications instead of decompositions to scale to big networks. On hard-to-optimize problems it beat Adam and other optimizers.

Abstract

Quasi-Newton (QN) methods have long been among the most effective methods for large-scale unconstrained convex optimization. Two obstacles have limited their use in deep learning: non-convexity and enormous parameter sizes. We introduce SoftServe, a family of QN methods designed to overcome these obstacles without line searches or ad hoc curvature corrections. SoftServe derives positivedefinite curvature estimates from the variational objective of Berglund et al. (2025), even in the presence of negative curvature. We develop diagonal and Kroneckerfactored variants that preserve positive definiteness by construction and scale to massive neural networks. Finally, SoftServe relies on the stable coupled Newton-Schulz iteration for the required matrix operations, replacing costly matrix decompositions with GPU-friendly matrix multiplications. SoftServe excels on problems that are severely ill-conditioned, including tasks such as recurrent networks, deep autoencoders, physics-informed neural networks, and a 136M-parameter physics-informed diffusion model, often achieving lower losses than established baselines including Adam, Muon, and SOAP.

Joohwan Ko, Tetiana Parshakova, Diana Cai, Robert M. Gower
arXiv:2610.02182 · cs.LG, cs.AI · submitted Oct 1, 2026
abstract · pdf · html

add comment on HN