about
LoMA: Lossless Compressed Memory Attention (arxiv.org)
99 points by PaulHoule on Jan 27, 2024 | hide | past | pdf | 8 comments on HN

In plain words: The model is trained to squeeze its conversation memory into a shorter form every few words, keeping every detail instead of throwing some away as usual. It does this in one pass with no helper model, cutting memory and computing costs without loss.

Abstract

Large Language Models (LLMs) face limitations due to the high demand on GPU memory and computational resources when handling long contexts. While sparsify the Key-Value (KV) cache of transformer model is a typical strategy to alleviate resource usage, it unavoidably results in the loss of information. We introduce Lossless Compressed Memory Attention (LoMA), a novel approach that enables lossless compression of the KV cache, thereby reducing the memory and computational demands during autoregressive generation. LoMA incorporates a specialized training or fine-tuning precedure alongside an autoregressive generation algorithm optimized for the compressed context. Our method compresses the KV cache after every $tc$ generated tokens with a compression ratio of $c$ and a target compressed length $t$, and this process occurs within a single inference pass without dependency on auxiliary models. We engineered an efficient training scheme involving specific inputs, attention masks, and position identifiers to instill this compression capability. Experimental validation has demonstrated that LoMA significantly reducing computational consumption and memory usage through achieving lossless KV cache compression.

Yumeng Wang, Zhenyang Xiao
arXiv:2401.09486 · cs.LG, cs.CL · submitted Jan 16, 2024 · updated Feb 4, 2024
abstract · pdf · html

add comment on HN

Cool, but poorly written. Why on page 4 spend half of it on standard attention (which should be assumed knowledge to a reader at this point) but then not explain equation 10? what is L? Why is there an identity matrix? I don't need equations 6-9 but I sure do need more information on 10-14. I hope there's code

What a weird line in Figure 2

> Note: L_LM represents L_LM, and L_Repeat represents L_Repeat, and Loss represents L.

Tautology is not helpful here.

And what's everyone's aversion to log plots? Figure 3 is unreadable but would be perfect with log.

And where's the Appendix? It's referenced from the paper.... I'm also unconvinced it's lossless

They spend that much time reiterating well-trodden established information because that's the easiest to redundantly speak to. All of the novel stuff is suspiciously vague and detail is left to the imagination. Because the author is fudging things.
This is actually something I really hate about ML research and why I get upset when people say "don't need math." I see a lot of unnecessary equations in papers (dear god, we don't have to write attention in every ViT paper or the density function in every diffusion paper) but then no math around where it matters. It feels like a strategy to get people to just turn on autopilot, since that's what you'll naturally do. It's also weird when you get reviewers asking for this. I've had a reviewer reject a work, saying that I should include the equation for attention... in 2023... There's a reason I don't value conferences and why I tell people they aren't good signals for quality. Hell, we're doing more peer review here than most papers get.

I'm not unconvinced the work isn't novel (but novelty is a function of reader's background knowledge) but I think their claims are too strong. Were I the reviewer I would reject for sure based on the unsubstantiated claim of lossless alone, without explanation. Even if a work was great and useful (though conferences make this difficult to make such types of fixes). The method does seem interesting and there is evidence for its utility, but I do think it needs a lot of additional clarity and support of the claims. In fact, I'll go ahead and say what's happening here is how I'd love to see modern peer review work, because we should be updating works and collaborating like this. Which isn't just shitting on the paper. It's constructive.

Edit: Sorry, got a bit ranty. CVPR reviews are out. I'll just say I'm so glad they made a big deal about LLMs to "protect the integrity of reviews" because... LLMs were the problems....

Indeed, this is also one of the worst abstracts to a paper ever. 90% of it is a generic introduction, and the rest goes out of its way to establish any quantitative results.
Very interesting, but on my initial skim through, I don’t understand how this technique is lossless? Reproduced from the methods section:

1. Select a sequence of tc tokens that the model has already generated or completed predicting as the reading area. 2. Insert t '<m>' tokens at once after the reading area to serve as the memory area. 3. The model performs a single inference on the memory area, but discards the model's output, retaining only the KV pairs from each layer. 4. Discard the reading area, and the model continues generating text from after the memory area.

Isn’t the memory area a lossy compression of the reading area?

The paper is very confusing and should have a title change. In their results (4.2.1) they say

> The observation that L_Repeat converges rapidly to a value close to zero under smaller compression ratios is significant. It demonstrates that the method is highly effective in compressing information losslessly into memory tokens

So I'm also unconvinced

Given these responses, should I use HackerNews as a pre-submission review process?!
Depends, are you looking to have a good paper or are you looking to get published in a venue? Because I'll say that you'll get better peer review anywhere (HN, Twitter, Reddit, etc) than you'll get in a conference.

So... yes?