about
xAI: A Spectral Condition for Feature Learning (arxiv.org)
3 points by georgehill on Dec 19, 2023 | hide | past | pdf | discuss on HN

In plain words: Sizing a layer's weights and updates by the square root of outputs over inputs, using the largest stretch the matrix applies to inputs, keeps representations changing at any width. It simply derives the popular maximal-update recipe, unlike the usual tweaks to individual weight entries.

Abstract · A Spectral Condition for Feature Learning

The push to train ever larger neural networks has motivated the study of initialization and training at large network width. A key challenge is to scale training so that a network's internal representations evolve nontrivially at all widths, a process known as feature learning. Here, we show that feature learning is achieved by scaling the spectral norm of weight matrices and their updates like $\sqrt{\texttt{fan-out}/\texttt{fan-in}}$, in contrast to widely used but heuristic scalings based on Frobenius norm and entry size. Our spectral scaling analysis also leads to an elementary derivation of \emph{maximal update parametrization}. All in all, we aim to provide the reader with a solid conceptual understanding of feature learning in neural networks.

Greg Yang, James B. Simon, Jeremy Bernstein
arXiv:2310.17813 · cs.LG · submitted Oct 26, 2023 · updated May 14, 2024
abstract · pdf · html

add comment on HN