about
Putting an End to End-to-End: Gradient-Isolated Learning of Representations (arxiv.org)
2 points by 1_over_n on Mar 3, 2020 | hide | past | pdf | discuss on HN

In plain words: A deep network is split into stacked modules, each trained on its own to keep as much information as possible from its input, with no labels and no global error signal sent backward. The top module's features nearly matched fully end-to-end training on audio and image classification, while modules could be trained separately in parallel.

Abstract · Putting An End to End-to-End: Gradient-Isolated Learning of Representations

We propose a novel deep learning method for local self-supervised representation learning that does not require labels nor end-to-end backpropagation but exploits the natural order in data instead. Inspired by the observation that biological neural networks appear to learn without backpropagating a global error signal, we split a deep neural network into a stack of gradient-isolated modules. Each module is trained to maximally preserve the information of its inputs using the InfoNCE bound from Oord et al. [2018]. Despite this greedy training, we demonstrate that each module improves upon the output of its predecessor, and that the representations created by the top module yield highly competitive results on downstream classification tasks in the audio and visual domain. The proposal enables optimizing modules asynchronously, allowing large-scale distributed training of very deep neural networks on unlabelled datasets.

Sindy Löwe, Peter O'Connor, Bastiaan S. Veeling
arXiv:1905.11786 · cs.LG, cs.AI, stat.ML · submitted May 28, 2019 · updated Jan 27, 2020
abstract · pdf · html · Honorable Mention for Outstanding New Directions Paper Award at NeurIPS 2019

add comment on HN