about
Self-supervised learning through the eyes of a child (arxiv.org)
1 point by spekcular on Dec 8, 2020 | hide | past | pdf | discuss on HN

In plain words: A vision system learned from video filmed by cameras worn by three young children, finding patterns in raw footage with no labels. It built strong high-level visual concepts, unlike the usual training on huge labeled image sets, suggesting generic learning alone can explain much early knowledge.

Abstract

Within months of birth, children develop meaningful expectations about the world around them. How much of this early knowledge can be explained through generic learning mechanisms applied to sensory data, and how much of it requires more substantive innate inductive biases? Addressing this fundamental question in its full generality is currently infeasible, but we can hope to make real progress in more narrowly defined domains, such as the development of high-level visual categories, thanks to improvements in data collecting technology and recent progress in deep learning. In this paper, our goal is precisely to achieve such progress by utilizing modern self-supervised deep learning methods and a recent longitudinal, egocentric video dataset recorded from the perspective of three young children (Sullivan et al., 2020). Our results demonstrate the emergence of powerful, high-level visual representations from developmentally realistic natural videos using generic self-supervised learning objectives.

A. Emin Orhan, Vaibhav V. Gupta, Brenden M. Lake
arXiv:2007.16189 · cs.CV, cs.LG, cs.NE · submitted Jul 31, 2020 · updated Dec 15, 2020
abstract · pdf · html · Published as a conference paper at NeurIPS 2020; v3 adds a reference, fixes a typo

add comment on HN