In plain words: An unsupervised training trick learns image features by predicting what comes next in a picture from what came before, capturing patterns in natural images without any labels. With these features, classifiers need 2-5x fewer labeled examples than ones trained straight on raw pixels.
Abstract
Human observers can learn to recognize new categories of images from a handful of examples, yet doing so with artificial ones remains an open challenge. We hypothesize that data-efficient recognition is enabled by representations which make the variability in natural signals more predictable. We therefore revisit and improve Contrastive Predictive Coding, an unsupervised objective for learning such representations. This new implementation produces features which support state-of-the-art linear classification accuracy on the ImageNet dataset. When used as input for non-linear classification with deep neural networks, this representation allows us to use 2-5x less labels than classifiers trained directly on image pixels. Finally, this unsupervised representation substantially improves transfer learning to object detection on the PASCAL VOC dataset, surpassing fully supervised pre-trained ImageNet classifiers.
Olivier J. Hénaff, Aravind Srinivas, Jeffrey De Fauw, Ali Razavi, Carl Doersch, S. M. Ali Eslami, Aaron van den Oord
arXiv:1905.09272 · cs.CV, cs.LG · submitted May 22, 2019 · updated Jul 1, 2020
abstract · pdf · html
Interesting paper on an unsupervised pre-training approach using overlapping image fields to massively reduce the amount of labeled data required to train image classification models.
Reminds me a lot of the way word2vec is trained with overlapping word vectors and negative sampling...