In plain words: A vision system that turns raw pixels into explicit object-and-relation structures, so its inner representations spell out how things relate instead of just bundling features. Trained step by step on simple visual reasoning puzzles, it then learns brand-new tasks more easily than standard networks.
Abstract
With a view to bridging the gap between deep learning and symbolic AI, we present a novel end-to-end neural network architecture that learns to form propositional representations with an explicitly relational structure from raw pixel data. In order to evaluate and analyse the architecture, we introduce a family of simple visual relational reasoning tasks of varying complexity. We show that the proposed architecture, when pre-trained on a curriculum of such tasks, learns to generate reusable representations that better facilitate subsequent learning on previously unseen tasks when compared to a number of baseline architectures. The workings of a successfully trained model are visualised to shed some light on how the architecture functions.
Murray Shanahan, Kyriacos Nikiforou, Antonia Creswell, Christos Kaplanis, David Barrett, Marta Garnelo
arXiv:1905.10307 · cs.LG, stat.ML · submitted May 24, 2019 · updated Jun 23, 2020
abstract · pdf · html · In Proceedings ICML 2020
This paper's section on related work calls out a recent paper by Asai
> Asai [1], whose paper was published while the present work was in progress, describes an architecture with some similarities to the PrediNet, but also some notable differences. For example, Asai’s architecture assumes an input representation in symbolic form where the objects have already been segmented. By contrast, in the present architecture, the input CNN and the PrediNet’s dot-product attention mechanism together learn what constitutes an object.
re: [1], Asai's paper: https://arxiv.org/abs/1902.08093 .
Asai has released code for earlier papers on github: "LatPlan : A domain-independent, image-based classical planner". https://github.com/guicho271828/latplan
it reads as if the code for [1] may be open-sourced in future "Asai, M.: 2019. Unsupervised Grounding of Plannable First-Order Logic Representation from Images (code not yet available) "
[+] research-quality, i.e., in whatever random state the prototype code it happens to be in, without any guarantees about it compiling, or running, or behaving correctly, having an automated test suite, or whatever. in all of these cases, it is _technically_ trivial to make it available to the wider research community