about
Learning Curve Theory (arxiv.org)
2 points by Anon84 on Feb 28, 2021 | hide | past | pdf | discuss on HN

In plain words: A simple toy model makes test error shrink as a power of the training set size, and is analyzed to see if that power depends on the data's shape. Unlike usual models, which only fall at n^-1/2 or n^-1, it can show any rate.

Abstract

Recently a number of empirical "universal" scaling law papers have been published, most notably by OpenAI. `Scaling laws' refers to power-law decreases of training or test error w.r.t. more data, larger neural networks, and/or more compute. In this work we focus on scaling w.r.t. data size $n$. Theoretical understanding of this phenomenon is largely lacking, except in finite-dimensional models for which error typically decreases with $n^{-1/2}$ or $n^{-1}$, where $n$ is the sample size. We develop and theoretically analyse the simplest possible (toy) model that can exhibit $n^{-β}$ learning curves for arbitrary power $β>0$, and determine whether power laws are universal or depend on the data distribution.

Marcus Hutter
arXiv:2102.04074 · cs.LG, stat.ML · submitted Feb 8, 2021
abstract · pdf · html · 26 pages, 6 Figures

add comment on HN
Also discussed: Feb 2021 (2 points, 0 comments)