about
The Extractor: a drop-in replacement for self attention in Transformers (arxiv.org)
1 point by groar on Aug 16, 2023 | hide | past | pdf | discuss on HN

In plain words: They replace the Transformer's self-attention step—the part where every word compares with every other word—with simpler "extractors" that have much shorter calculation chains. The strongest version does better than normal self-attention, while lighter versions match or beat it using less computing power and memory.

Abstract · Attention Is Not All You Need Anymore

In recent years, the popular Transformer architecture has achieved great success in many application areas, including natural language processing and computer vision. Many existing works aim to reduce the computational and memory complexity of the self-attention mechanism in the Transformer by trading off performance. However, performance is key for the continuing success of the Transformer. In this paper, a family of drop-in replacements for the self-attention mechanism in the Transformer, called the Extractors, is proposed. Four types of the Extractors, namely the super high-performance Extractor (SHE), the higher-performance Extractor (HE), the worthwhile Extractor (WE), and the minimalist Extractor (ME), are proposed as examples. Experimental results show that replacing the self-attention mechanism with the SHE evidently improves the performance of the Transformer, whereas the simplified versions of the SHE, i.e., the HE, the WE, and the ME, perform close to or better than the self-attention mechanism with less computational and memory complexity. Furthermore, the proposed Extractors have the potential or are able to run faster than the self-attention mechanism since their critical paths of computation are much shorter. Additionally, the sequence prediction problem in the context of text generation is formulated using variable-length discrete-time Markov chains, and the Transformer is reviewed based on our understanding.

Zhe Chen
arXiv:2308.07661 · cs.LG, cs.CL, cs.NE · submitted Aug 15, 2023 · updated Sep 19, 2023
abstract · pdf · html

add comment on HN