about
A Neural Scaling Law from the Dimension of the Data Manifold (arxiv.org)
2 points by noanabeshima on Nov 13, 2020 | hide | past | pdf | discuss on HN

In plain words: Bigger neural networks make fewer mistakes at a predictable rate, and this study shows the rate's steepness is set by how many independent directions the data has: error falls about 4 divided by that number. Tests on controlled, image, and language data confirmed it.

Abstract

When data is plentiful, the loss achieved by well-trained neural networks scales as a power-law $L \propto N^{-α}$ in the number of network parameters $N$. This empirical scaling law holds for a wide variety of data modalities, and may persist over many orders of magnitude. The scaling law can be explained if neural models are effectively just performing regression on a data manifold of intrinsic dimension $d$. This simple theory predicts that the scaling exponents $α\approx 4/d$ for cross-entropy and mean-squared error losses. We confirm the theory by independently measuring the intrinsic dimension and the scaling exponents in a teacher/student framework, where we can study a variety of $d$ and $α$ by dialing the properties of random teacher networks. We also test the theory with CNN image classifiers on several datasets and with GPT-type language models.

Utkarsh Sharma, Jared Kaplan
arXiv:2004.10802 · cs.LG, stat.ML · submitted Apr 22, 2020
abstract · pdf · html · 16+12 pages, 11+11 figures

add comment on HN