about
Fast Scene Understanding – Hinton's Latest Paper (arxiv.org)
2 points by p1esk on Mar 31, 2016 | hide | past | pdf | discuss on HN

In plain words: A network spots one object at a time and decides how many objects to examine, letting it break a scene into parts without being shown any labels. It produced as accurate inferences as supervised versions given labels, and generalized better to new scenes.

Abstract · Attend, Infer, Repeat: Fast Scene Understanding with Generative Models

We present a framework for efficient inference in structured image models that explicitly reason about objects. We achieve this by performing probabilistic inference using a recurrent neural network that attends to scene elements and processes them one at a time. Crucially, the model itself learns to choose the appropriate number of inference steps. We use this scheme to learn to perform inference in partially specified 2D models (variable-sized variational auto-encoders) and fully specified 3D models (probabilistic renderers). We show that such models learn to identify multiple objects - counting, locating and classifying the elements of a scene - without any supervision, e.g., decomposing 3D images with various numbers of objects in a single forward pass of a neural network. We further show that the networks produce accurate inferences when compared to supervised counterparts, and that their structure leads to improved generalization.

S. M. Ali Eslami, Nicolas Heess, Theophane Weber, Yuval Tassa, David Szepesvari, Koray Kavukcuoglu, Geoffrey E. Hinton
arXiv:1603.08575 · cs.CV, cs.LG · submitted Mar 28, 2016 · updated Aug 12, 2016
abstract · pdf · html

add comment on HN