about
Survey of Dropout Methods for Deep Neural Networks (arxiv.org)
88 points by Anon84 on May 1, 2019 | hide | past | pdf | 5 comments on HN

In plain words: Dropout randomly switches off parts of a neural network during training so it cannot lean too heavily on any one path. This survey traces its history and uses — preventing overfitting, shrinking models, gauging certainty — and finds it now extends to image and sequence layers.

Abstract

Dropout methods are a family of stochastic techniques used in neural network training or inference that have generated significant research interest and are widely used in practice. They have been successfully applied in neural network regularization, model compression, and in measuring the uncertainty of neural network outputs. While original formulated for dense neural network layers, recent advances have made dropout methods also applicable to convolutional and recurrent neural network layers. This paper summarizes the history of dropout methods, their various applications, and current areas of research interest. Important proposed methods are described in additional detail.

Alex Labach, Hojjat Salehinejad, Shahrokh Valaee
arXiv:1904.13310 · cs.NE, cs.AI, cs.LG · submitted Apr 25, 2019 · updated Oct 25, 2019
abstract · pdf · html

add comment on HN

When in doubt, use random dropouts
Is drop out still empirical or are there any proof of why it works in the overall model?

I recall reading up on CNN and playing around with it and it was interesting to add random drop off in there but was never explained why it works. I think the general thinking of why it works is that the network is overfitting so randomly dropping node is required for generalization?

Addressing your second question. Informally, dropping nodes fights overfitting by creating subsample architectures of which are essentially thinned out networks of the one you've designed. Having trained on these sub nets means you've effectively combined the learning of a few different models and in doing so have generalized beyond the capabilities of your original "single" architecture.
My understanding is that it avoids overfitting when data points are highly correlated.

For example, if you use image augmentation to generate additional data, your augmented images are going to be highly correlated to their parent image leading to overfitting of the data. By using random dropout, this overfitting can be somewhat mitigated.

For linear models it can be shown to be equivalent to weight decay. For nonlinear ones, it empirically behaves as a regularizer.,