about
Reformer: The Efficient Transformer (arxiv.org)
1 point by blopeur on Jan 19, 2020 | hide | past | pdf | discuss on HN

In plain words: Instead of letting every word compare with every other word, it sorts words into buckets by a quick fingerprint so each checks only a few similar ones. It matches normal Transformers' quality while using much less memory and running faster on long texts.

Abstract

Large Transformer models routinely achieve state-of-the-art results on a number of tasks but training these models can be prohibitively costly, especially on long sequences. We introduce two techniques to improve the efficiency of Transformers. For one, we replace dot-product attention by one that uses locality-sensitive hashing, changing its complexity from O($L^2$) to O($L\log L$), where $L$ is the length of the sequence. Furthermore, we use reversible residual layers instead of the standard residuals, which allows storing activations only once in the training process instead of $N$ times, where $N$ is the number of layers. The resulting model, the Reformer, performs on par with Transformer models while being much more memory-efficient and much faster on long sequences.

Nikita Kitaev, Łukasz Kaiser, Anselm Levskaya
arXiv:2001.04451 · cs.LG, cs.CL, stat.ML · submitted Jan 13, 2020 · updated Feb 18, 2020
abstract · pdf · html · ICLR 2020

add comment on HN