In plain words: A learning approach finds hidden structure in unlabeled data by building a compact summary that explains as much shared redundancy as possible. Unlike usual training on labeled examples, it needed no labels and succeeded on tasks in human behavior, biology, and language.
Abstract
Learning by children and animals occurs effortlessly and largely without obvious supervision. Successes in automating supervised learning have not translated to the more ambiguous realm of unsupervised learning where goals and labels are not provided. Barlow (1961) suggested that the signal that brains leverage for unsupervised learning is dependence, or redundancy, in the sensory environment. Dependence can be characterized using the information-theoretic multivariate mutual information measure called total correlation. The principle of Total Cor-relation Ex-planation (CorEx) is to learn representations of data that "explain" as much dependence in the data as possible. We review some manifestations of this principle along with successes in unsupervised learning problems across diverse domains including human behavior, biology, and language.
Greg Ver Steeg
arXiv:1706.08984 · stat.ML · submitted Jun 27, 2017
abstract · pdf · html · Invited contribution for IJCAI 2017 Early Career Spotlight. 5 pages, 1 figure