In plain words: A network learns visual features without labels by repeatedly grouping its own outputs and using those groups as labels to retrain itself. On large image collections it beat the best previous unlabeled training approach by a wide margin on standard tests.
Abstract · Deep Clustering for Unsupervised Learning of Visual Features
Clustering is a class of unsupervised learning methods that has been extensively applied and studied in computer vision. Little work has been done to adapt it to the end-to-end training of visual features on large scale datasets. In this work, we present DeepCluster, a clustering method that jointly learns the parameters of a neural network and the cluster assignments of the resulting features. DeepCluster iteratively groups the features with a standard clustering algorithm, k-means, and uses the subsequent assignments as supervision to update the weights of the network. We apply DeepCluster to the unsupervised training of convolutional neural networks on large datasets like ImageNet and YFCC100M. The resulting model outperforms the current state of the art by a significant margin on all the standard benchmarks.
Mathilde Caron, Piotr Bojanowski, Armand Joulin, Matthijs Douze
arXiv:1807.05520 · cs.CV · submitted Jul 15, 2018 · updated Mar 18, 2019
abstract · pdf · html · Accepted at ECCV 2018