about
Identity Mappings in Deep Residual Networks (arxiv.org)
3 points by bluelinespecial on Mar 17, 2016 | hide | past | pdf | discuss on HN

In plain words: By making each skip connection a plain copy and moving the activation after the addition, signals pass unchanged between any two blocks, so very deep networks train more easily. A 1001-layer version reached 4.62% error on CIFAR-10, beating the original residual design.

Abstract

Deep residual networks have emerged as a family of extremely deep architectures showing compelling accuracy and nice convergence behaviors. In this paper, we analyze the propagation formulations behind the residual building blocks, which suggest that the forward and backward signals can be directly propagated from one block to any other block, when using identity mappings as the skip connections and after-addition activation. A series of ablation experiments support the importance of these identity mappings. This motivates us to propose a new residual unit, which makes training easier and improves generalization. We report improved results using a 1001-layer ResNet on CIFAR-10 (4.62% error) and CIFAR-100, and a 200-layer ResNet on ImageNet. Code is available at: https://github.com/KaimingHe/resnet-1k-layers

Kaiming He, Xiangyu Zhang, Shaoqing Ren, Jian Sun
arXiv:1603.05027 · cs.CV, cs.LG · submitted Mar 16, 2016 · updated Jul 25, 2016
abstract · pdf · html · ECCV 2016 camera-ready

add comment on HN