In plain words: Instead of reading every input token, it stores them in a search index and pulls only the few most relevant at each step, letting a pretrained summarizer handle unlimited text without retraining. It summarized 500,000-token books without cutting anything, beating the usual truncation approach.
Abstract
Since the proposal of transformers, these models have been limited to bounded input lengths, because of their need to attend to every token in the input. In this work, we propose Unlimiformer: a general approach that wraps any existing pretrained encoder-decoder transformer, and offloads the cross-attention computation to a single k-nearest-neighbor (kNN) index, while the returned kNN distances are the attention dot-product scores. This kNN index can be kept on either the GPU or CPU memory and queried in sub-linear time; this way, we can index practically unlimited input sequences, while every attention head in every decoder layer retrieves its top-k keys, instead of attending to every key. We evaluate Unlimiformer on several long-document and book-summarization benchmarks, showing that it can process even 500k token-long inputs from the BookSum dataset, without any input truncation at test time. We demonstrate that Unlimiformer improves pretrained models such as BART and Longformer by extending them to unlimited inputs without additional learned weights and without modifying their code. We make our code and models publicly available at https://github.com/abertsch72/unlimiformer .
Amanda Bertsch, Uri Alon, Graham Neubig, Matthew R. Gormley
arXiv:2305.01625 · cs.CL · submitted May 2, 2023 · updated Oct 30, 2023
abstract · pdf · html · NeurIPS 2023
2. This idea is quite similar to retrieval transformers and Hopfield networks which have been known and published for several years now. It's not really that novel.
3. Due to the preceding points, the title can easily mislead people. It's not really a conventional transformer, and it's not a breakthrough.
4. This paper is a preprint and not peer-reviewed.
"I generally don't enjoy seeing preprints like this going to the top of Hacker News. This would be a higher quality submission if the paper was peer-reviewed or put into a greater context, like a blog post discussion or something like that."
Let me retract this and say something a bit nicer :) I personally think there this specific preprint making it to the top of HN is potentially harmful, because of the hype around LLMs, the diverse audience of readers here, and the specific title that implies a claim of "transformer with unlimited context length", when this is misleading. I don't have anything against preprints in general - a lot of work outside of the peer-review process ends up being very impactful.