about
The intriguing role of module criticality in the generalization of deep networks (arxiv.org)
1 point by scribu on Dec 4, 2019 | hide | past | pdf | discuss on HN

In plain words: Some parts of a trained network matter more: resetting them to their starting values sharply hurts accuracy. Scoring each part by the shape of the error dip between its start and final values predicts which architectures generalize better, unlike earlier scores.

Abstract

We study the phenomenon that some modules of deep neural networks (DNNs) are more critical than others. Meaning that rewinding their parameter values back to initialization, while keeping other modules fixed at the trained parameters, results in a large drop in the network's performance. Our analysis reveals interesting properties of the loss landscape which leads us to propose a complexity measure, called module criticality, based on the shape of the valleys that connects the initial and final values of the module parameters. We formulate how generalization relates to the module criticality, and show that this measure is able to explain the superior generalization performance of some architectures over others, whereas earlier measures fail to do so.

Niladri S. Chatterji, Behnam Neyshabur, Hanie Sedghi
arXiv:1912.00528 · cs.LG, stat.ML · submitted Dec 2, 2019 · updated Feb 14, 2020
abstract · pdf · html

add comment on HN