about
Recurrent Models of Visual Attention, DeepMind (arxiv.org)
3 points by albertzeyer on Dec 18, 2014 | hide | past | pdf | discuss on HN

In plain words: Instead of processing every pixel, this system looks at one small patch at a time, choosing where to look next. It beat a standard whole-image network on cluttered pictures and learned to track a moving object without being told to.

Abstract · Recurrent Models of Visual Attention

Applying convolutional neural networks to large images is computationally expensive because the amount of computation scales linearly with the number of image pixels. We present a novel recurrent neural network model that is capable of extracting information from an image or video by adaptively selecting a sequence of regions or locations and only processing the selected regions at high resolution. Like convolutional neural networks, the proposed model has a degree of translation invariance built-in, but the amount of computation it performs can be controlled independently of the input image size. While the model is non-differentiable, it can be trained using reinforcement learning methods to learn task-specific policies. We evaluate our model on several image classification tasks, where it significantly outperforms a convolutional neural network baseline on cluttered images, and on a dynamic visual control problem, where it learns to track a simple object without an explicit training signal for doing so.

Volodymyr Mnih, Nicolas Heess, Alex Graves, Koray Kavukcuoglu
arXiv:1406.6247 · cs.LG, cs.CV, stat.ML · submitted Jun 24, 2014
abstract · pdf · html

add comment on HN