about
Object-based attention for spatio-temporal reasoning (arxiv.org)
2 points by jedharris on Dec 16, 2020 | hide | past | pdf | 2 comments on HN

In plain words: The network learns object-like pieces from images, lets each piece pay attention to the others, and practices predicting how they change over time without labels. It beat task-specific modular systems built for each problem on all three reasoning tasks.

Abstract · Attention over learned object embeddings enables complex visual reasoning

Neural networks have achieved success in a wide array of perceptual tasks but often fail at tasks involving both perception and higher-level reasoning. On these more challenging tasks, bespoke approaches (such as modular symbolic components, independent dynamics models or semantic parsers) targeted towards that specific type of task have typically performed better. The downside to these targeted approaches, however, is that they can be more brittle than general-purpose neural networks, requiring significant modification or even redesign according to the particular task at hand. Here, we propose a more general neural-network-based approach to dynamic visual reasoning problems that obtains state-of-the-art performance on three different domains, in each case outperforming bespoke modular approaches tailored specifically to the task. Our method relies on learned object-centric representations, self-attention and self-supervised dynamics learning, and all three elements together are required for strong performance to emerge. The success of this combination suggests that there may be no need to trade off flexibility for performance on problems involving spatio-temporal or causal-style reasoning. With the right soft biases and learning objectives in a neural network we may be able to attain the best of both worlds.

David Ding, Felix Hill, Adam Santoro, Malcolm Reynolds, Matt Botvinick
arXiv:2012.08508 · cs.CV, cs.AI, cs.CL, cs.LG · submitted Dec 15, 2020 · updated Oct 26, 2021
abstract · pdf · html · 22 pages, 5 figures

add comment on HN

This paper drives a stake through the heart of claims that explicit symbolic reasoning is needed for intelligent systems. It builds extra-symbolic (deep learning) systems that substantially outperform symbolic reasoning on tests specifically designed to require symbolic reasoning.

Since today's programming is 1000% explicitly symbolic, the potential shift to better methods that are extra-symbolic will be pretty disruptive!

This paper is a great overview of the state of the art in building flexible generalizable deep learning systems. Not easy to read, but if you puzzle through it you'll understand where the field is and where it is going.