about
Seer: Self-Supervised Pretraining of Visual Features in the Wild (arxiv.org)
1 point by ArtWomb on Mar 8, 2021 | hide | past | pdf | discuss on HN

In plain words: A giant vision model learns by training on 1 billion random, unlabeled web photos instead of a carefully curated labeled set. It reached 84.2% accuracy on ImageNet, beating the best previous label-free model by 1%, and stayed strong with only 10% of the labels.

Abstract · Self-supervised Pretraining of Visual Features in the Wild

Recently, self-supervised learning methods like MoCo, SimCLR, BYOL and SwAV have reduced the gap with supervised methods. These results have been achieved in a control environment, that is the highly curated ImageNet dataset. However, the premise of self-supervised learning is that it can learn from any random image and from any unbounded dataset. In this work, we explore if self-supervision lives to its expectation by training large models on random, uncurated images with no supervision. Our final SElf-supERvised (SEER) model, a RegNetY with 1.3B parameters trained on 1B random images with 512 GPUs achieves 84.2% top-1 accuracy, surpassing the best self-supervised pretrained model by 1% and confirming that self-supervised learning works in a real world setting. Interestingly, we also observe that self-supervised models are good few-shot learners achieving 77.9% top-1 with access to only 10% of ImageNet. Code: https://github.com/facebookresearch/vissl

Priya Goyal, Mathilde Caron, Benjamin Lefaudeux, Min Xu, Pengchao Wang, Vivek Pai, Mannat Singh, Vitaliy Liptchinsky, Ishan Misra, Armand Joulin, Piotr Bojanowski
arXiv:2103.01988 · cs.CV, cs.AI · submitted Mar 2, 2021 · updated Mar 5, 2021
abstract · pdf · html

add comment on HN