In plain words: Instead of scoring pairs of words, this attention scores triples at once and multiplies its output vectors together to update each entity's representation. In reinforcement learning tasks, it gave agents a better handle on logical reasoning than standard dot-product attention.
Abstract · Logic and the $2$-Simplicial Transformer
We introduce the $2$-simplicial Transformer, an extension of the Transformer which includes a form of higher-dimensional attention generalising the dot-product attention, and uses this attention to update entity representations with tensor products of value vectors. We show that this architecture is a useful inductive bias for logical reasoning in the context of deep reinforcement learning.
James Clift, Dmitry Doryn, Daniel Murfet, James Wallbridge
arXiv:1909.00668 · cs.LG, cs.LO, stat.ML · submitted Sep 2, 2019
abstract · pdf · html