about
Hierarchical transformers are more efficient language models (arxiv.org)
91 points by beefman on Nov 4, 2021 | hide | past | pdf | 5 comments on HN

In plain words: The model shrinks the sequence of words it processes as it goes deeper, then grows it back before outputting, so long texts cost far less to handle. At the same computing budget it beat a standard transformer, matching its quality with less work.

Abstract · Hierarchical Transformers Are More Efficient Language Models

Transformer models yield impressive results on many NLP and sequence modeling tasks. Remarkably, Transformers can handle long sequences which allows them to produce long coherent outputs: full paragraphs produced by GPT-3 or well-structured images produced by DALL-E. These large language models are impressive but also very inefficient and costly, which limits their applications and accessibility. We postulate that having an explicit hierarchical architecture is the key to Transformers that efficiently handle long sequences. To verify this claim, we first study different ways to downsample and upsample activations in Transformers so as to make them hierarchical. We use the best performing upsampling and downsampling layers to create Hourglass - a hierarchical Transformer language model. Hourglass improves upon the Transformer baseline given the same amount of computation and can yield the same results as Transformers more efficiently. In particular, Hourglass sets new state-of-the-art for Transformer models on the ImageNet32 generation task and improves language modeling efficiency on the widely studied enwik8 benchmark.

Piotr Nawrot, Szymon Tworkowski, Michał Tyrolski, Łukasz Kaiser, Yuhuai Wu, Christian Szegedy, Henryk Michalewski
arXiv:2110.13711 · cs.LG, cs.CL · submitted Oct 26, 2021 · updated Apr 16, 2022
abstract · pdf · html

add comment on HN

This work is quite interesting, and I'm always happy to see such large improvements for memory usage and computation over existing high performant models. I imagine transformer-based models will be with the ML community for a very long time (like ResNets).

I will say that it is somewhat comical that an already existing paper titled "Hierarchical Transformers for Long Document Classification" isn't mentioned in the related work. But to be fair, the two papers are only modestly similar.

I agree that Transformers are here to stay. The basic building block (self-attention layers) seem to me like the "new fully-connected layer" - the natural way to connect layers and build a deep net. Except that with FC-layers the activations can only be fixed-sized vectors. But the self-attention, each layer can have a variable-sized bag of vectors, and you just need to encode their relationship to each other somehow. This is clearly successful for text using spectral positional encoding. It's starting to work for images, with 2D positional encoding. There's every reason to think it will work for many other data types.

It seems to me the key barrier is the high computational overhead for self-attention. But in highly-parallel vector-math world (GPUs, TPUs, NPUs, etc) Moore's law marches on, with little end in sight, because parallelism works great. That said, making them more efficient, like this paper, will certainly help their adoption.

They work for everything.
The models in the two papers are pretty different. If I were a reviewer I wouldn't expect TFA to cite that paper.
In their defense this field of study moves incredibly fast—there are hundreds of papers published every week. It would be a full-time job to simply keep up-to-date with the literature, to say nothing of doing your own research, so it’s hardly surprising the authors might not mention a given paper.