about
HyperAttention: Long-Context Attention in Near-Linear Time (arxiv.org)
73 points by kelseyfrog on Oct 10, 2023 | hide | past | pdf | 13 comments on HN

In plain words: Attention compares every word with every other word, so its cost grows with the square of the context length. HyperAttention uses hashing to spot the biggest comparisons and samples the rest, making ChatGLM2 inference 50% faster at 32k words with slightly worse quality.

Abstract · HyperAttention: Long-context Attention in Near-Linear Time

We present an approximate attention mechanism named HyperAttention to address the computational challenges posed by the growing complexity of long contexts used in Large Language Models (LLMs). Recent work suggests that in the worst-case scenario, quadratic time is necessary unless the entries of the attention matrix are bounded or the matrix has low stable rank. We introduce two parameters which measure: (1) the max column norm in the normalized attention matrix, and (2) the ratio of row norms in the unnormalized attention matrix after detecting and removing large entries. We use these fine-grained parameters to capture the hardness of the problem. Despite previous lower bounds, we are able to achieve a linear time sampling algorithm even when the matrix has unbounded entries or a large stable rank, provided the above parameters are small. HyperAttention features a modular design that easily accommodates integration of other fast low-level implementations, particularly FlashAttention. Empirically, employing Locality Sensitive Hashing (LSH) to identify large entries, HyperAttention outperforms existing methods, giving significant speed improvements compared to state-of-the-art solutions like FlashAttention. We validate the empirical performance of HyperAttention on a variety of different long-context length datasets. For example, HyperAttention makes the inference time of ChatGLM2 50\% faster on 32k context length while perplexity increases from 5.6 to 6.3. On larger context length, e.g., 131k, with causal masking, HyperAttention offers 5-fold speedup on a single attention layer.

Insu Han, Rajesh Jayaram, Amin Karbasi, Vahab Mirrokni, David P. Woodruff, Amir Zandieh
arXiv:2310.05869 · cs.LG, cs.AI · submitted Oct 9, 2023 · updated Dec 1, 2023
abstract · pdf · html

add comment on HN

"For example, HyperAttention makes the inference time of ChatGLM2 50% faster on 32k context length while perplexity increases from 5.6 to 6.3."

"when half of all attention layers are patched (i.e., 14 layers), we verify that most of the tasks do not degrade more than 13%."

According to the paper, for most tasks it reduces benchmark scores substantially. Perhaps to the point where a smaller model would yield better inference time and higher benchmarks.

However, summarization benchmarks see almost no degredation, great!

Smaller models will likely not have 32k context windows.
Why is that?
I assume it's because such large context takes lots of memory, so you might as well have smarter model if you are not gonna fit in small vram anyway
Personally, I have found that Mistral 7B (with its native 8K context, and decent results stretched out even more) is performing much better than llama 13B tunes for storytelling, where that long context is really important.

And I think the optimized backends should implement that sliding 16k context soon...

Anyway, point is a huge context really helps certain types of queries, and VRAM usage is reasonable with a 7B model.

ML researchers are playing scientists: tweak a few parameters in an LLM, re-train it on a largish dataset (need access to $$$ GPUs), find metrics on which the tweaked LLM makes a barely noticeable improvement and make the other metrics where it actually gets worse look insignificant, write a paper, upload to arxiv, and update your resume.
So true, if they were real scientists then they'd never publish negative results and only pass them along in their network to most efficiently gate keep the field!
> write a paper

Also publish yet another "SOTA framework" thats barebones and won't be maintained for very long!

The researchers who published the negative CFG LLM paper made an earnest effort to pull it into the popular frameworks. That really stands out in my memory, that is incredibly rare.

I'd rather the knowledge be out there, it helps you not go down dead end paths that you would otherwise have explored.
Every real ML position industry or academia explicitly asks for publications. The field gets what it asks for, and they get it good and hard.
I genuinely can't tell if this is supposed to be a criticism.

They tried something, it improved some things and made other things worse.

This paper presents formal results, apparently. Also, in the case of formal results, peer review by experts in the exact sub-field makes sense/increases trust, sure, but why not share on archive instead of waiting for a year (or however long the process takes in the particular instance)?
Worse are comparison papers where the researchers are evaluating models against a benchmark. In another world, this would have been a blog post.