about
Block-State Transformers (arxiv.org)
2 points by tosh on Dec 10, 2023 | hide | past | pdf | discuss on HN

In plain words: A new layer mixes a running summary of the whole text for long-range context with block-by-block attention for nearby words. It beat similar Transformer layers on language modeling and handled longer sequences, running over ten times faster than a block-recurrent Transformer layer when parallelized.

Abstract

State space models (SSMs) have shown impressive results on tasks that require modeling long-range dependencies and efficiently scale to long sequences owing to their subquadratic runtime complexity. Originally designed for continuous signals, SSMs have shown superior performance on a plethora of tasks, in vision and audio; however, SSMs still lag Transformer performance in Language Modeling tasks. In this work, we propose a hybrid layer named Block-State Transformer (BST), that internally combines an SSM sublayer for long-range contextualization, and a Block Transformer sublayer for short-term representation of sequences. We study three different, and completely parallelizable, variants that integrate SSMs and block-wise attention. We show that our model outperforms similar Transformer-based architectures on language modeling perplexity and generalizes to longer sequences. In addition, the Block-State Transformer demonstrates more than tenfold increase in speed at the layer level compared to the Block-Recurrent Transformer when model parallelization is employed.

Mahan Fathi, Jonathan Pilault, Orhan Firat, Christopher Pal, Pierre-Luc Bacon, Ross Goroshin
arXiv:2306.09539 · cs.CL, cs.LG · submitted Jun 15, 2023 · updated Oct 30, 2023
abstract · pdf · html · NeurIPS'23 - Thirty-seventh Conference on Neural Information Processing Systems

add comment on HN