In plain words: Inside a big randomly wired neural network, some subset of its connections already works as well as a target network, with no training at all. The proof shows this holds for any data and target, as long as the random network is big enough.
Abstract
The lottery ticket hypothesis (Frankle and Carbin, 2018), states that a randomly-initialized network contains a small subnetwork such that, when trained in isolation, can compete with the performance of the original network. We prove an even stronger hypothesis (as was also conjectured in Ramanujan et al., 2019), showing that for every bounded distribution and every target network with bounded weights, a sufficiently over-parameterized neural network with random weights contains a subnetwork with roughly the same accuracy as the target network, without any further training.
Eran Malach, Gilad Yehudai, Shai Shalev-Shwartz, Ohad Shamir
arXiv:2002.00585 · cs.LG, stat.ML · submitted Feb 3, 2020
abstract · pdf · html
I recommend for an overview:
- the original paper "The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks", https://arxiv.org/abs/1803.03635
- "Deconstructing Lottery Tickets: Zeros, Signs, and the Supermask" https://eng.uber.com/deconstructing-lottery-tickets/ showing that if we remove "non-winning tickets" before the training, the trained network still works well