about
Deep Networks with Stochastic Depth (arxiv.org)
70 points by nicklo on Apr 4, 2016 | hide | past | pdf | 8 comments on HN

In plain words: During training, randomly skip some layers so each batch runs through a shorter network, then use the full deep one at test time. This trains faster and cuts test error, letting networks pass 1200 layers and reach 4.91% error on CIFAR-10.

Abstract

Very deep convolutional networks with hundreds of layers have led to significant reductions in error on competitive benchmarks. Although the unmatched expressiveness of the many layers can be highly desirable at test time, training very deep networks comes with its own set of challenges. The gradients can vanish, the forward flow often diminishes, and the training time can be painfully slow. To address these problems, we propose stochastic depth, a training procedure that enables the seemingly contradictory setup to train short networks and use deep networks at test time. We start with very deep networks but during training, for each mini-batch, randomly drop a subset of layers and bypass them with the identity function. This simple approach complements the recent success of residual networks. It reduces training time substantially and improves the test error significantly on almost all data sets that we used for evaluation. With stochastic depth we can increase the depth of residual networks even beyond 1200 layers and still yield meaningful improvements in test error (4.91% on CIFAR-10).

Gao Huang, Yu Sun, Zhuang Liu, Daniel Sedra, Kilian Weinberger
arXiv:1603.09382 · cs.LG, cs.CV, cs.NE · submitted Mar 30, 2016 · updated Jul 28, 2016
abstract · pdf · html · first two authors contributed equally

add comment on HN
Also discussed: Apr 2016 (4 points, 0 comments) · Apr 2016 (3 points, 0 comments)

I suspect this was posted because of Delip Rao's write-up[1] (which I suggest might be a better link).

It's a nice - if somewhat controversial - summary.

40% speedup on DNN training with state-of-the-art results.

[1] http://deliprao.com/archives/134

It's kind of bonkers that this works. It suggests that the whole belief that layers are learning different representations is completely wrong: if layer three is expecting a certain kind of intermediate representation from layer two, and is then given the raw input, one would expect layer three to choke.

Instead, the depth seems to be giving something like a progressive unwinding of the feature space.

It would be interesting to compare the trained networks to networks trained in the usual way, to see if they're coming up with similar coefficients in spite of the different training methods, out if this is producing something completely different.

Note that this was done for 100-1000 layer depth. So each individual layer only slightly increases the "high-levelness" of the features. In the same sense, that is why Deep Residual network works - initially, all layers are close to identity.
Having not read the paper, something I find unclear: is it only the feedback path that is skipped, or is the feedforward path also skipped? The abstract mentions replacing the layer with an identity function. I'm not sure how this would work, wouldn't it change the result (i.e the encoding used by the following layer would be corrupted) if you just multiply the inputs by 1 and add them?

Otherwise, how precisely do you "skip" a layer without corrupting the training of lower layers?

Edit: the answer is in the definition of "skip layers", introduced in a previous paper: http://arxiv.org/abs/1512.03385 which introduces identity functions into the layer equation.. I guess I have more reading to do on this topic.

The full layer is skipped. I.e. replaced by identity. Sth like g(x) = switch(prob, x, f(x)).

Deep Residual Networks are similar but different. There, you add with identity. Sth like g(x) = x + f(x).

Yes, but my question was more, when the layer is "skipped", what happens to the input for the next layer? But clearly, it is designed such that the identity function still provides somehow useful information to the next layer. (i.e doesn't significantly transform its domain and range) I was wondering how this could work. It's just that, intuitively, I would think that the next layer is being trained on a specific transformation performed by the skipped layer, so I still don't fully understand how replacing a whole layer with the identity function doesn't completely mess up the training of all subsequent layers. But maybe the secret is that it only lasts for a small number of iterations, and perhaps this short-lived deviation actually helps inject some minima-escaping trajectory. (I have read that injecting random noise can have similar effects. Is this just a different kind of random noise?)
I'm reading Delip's followup post[1] and it reminds me how much of ANN stuff is till pretty much alchemy.

[1] http://deliprao.com/archives/137

This is literally one of the most exciting papers I have read recently that will have quite some impact on deep learning models. The major drawback of deep architectures today is training time and any.improvement to that will have a drastic effect on my productivity.

Right now I basically run N architectures on N GPUs at the same time to speed things up. And that's a luxury.