about
Fast KV Compaction via Attention Matching (arxiv.org)
73 points by cbracketdash 225 days ago | hide | past | pdf | 15 comments on HN

In plain words: Instead of shrinking a model's stored memory by summarizing the text, this builds tiny stand-in entries that reproduce the full context's attention outputs. It cuts the cache up to 50 times in seconds with little quality loss, unlike the slow training-based approach it replaces.

Abstract

Scaling language models to long contexts is often bottlenecked by the size of the key-value (KV) cache. In deployed settings, long contexts are typically managed through compaction in token space via summarization. However, summarization can be highly lossy, substantially harming downstream performance. Recent work on Cartridges has shown that it is possible to train highly compact KV caches in latent space that closely match full-context performance, but at the cost of slow and expensive end-to-end optimization. This work describes an approach for fast context compaction in latent space through Attention Matching, which constructs compact keys and values to reproduce attention outputs and preserve attention mass at a per-KV-head level. We show that this formulation naturally decomposes into simple subproblems, some of which admit efficient closed-form solutions. Within this framework, we develop a family of methods that significantly push the Pareto frontier of compaction time versus quality, achieving up to 50x compaction in seconds on some datasets with little quality loss.

Adam Zweiger, Xinghong Fu, Han Guo, Yoon Kim
arXiv:2602.16284 · cs.LG · submitted Feb 18, 2026 · updated May 26, 2026
abstract · pdf · html

add comment on HN

Considering the insanity of the AI arms race going on now, and the incredible sums of money be thrown at any slight advantage, is there any reason to believe that any meaningful AI breakthrough would be openly published for anyone to leverage?
These folks are MIT, so citations are valuable to them. Citations convert into prestige, academic career progression, or a favorable exit from academia into industry.

Also, I don't see why you couldn't patent this if you wanted to monetize it.

> Also, I don't see why you couldn't patent this if you wanted to monetize it.

We all just saw the prior art published for the public. That will preclude patenting this work. Further reduction to practice is required.

(I am not a lawyer).

Yes there is. Lots of researchers are more interested in making a contribution to societal flourishing than in making incredible sums of money. That’s why there’s still lots of top AI researchers in academia.
I do sometimes wonder -- if the transformers paper wasn't published, what would the industry be like? Would the same ideas have been put together in almost the same way weeks or months later somewhere else?
I would say yes.

The reality is that the money being thrown = the time of humans. I guess compute as well, but in terms of people doing innovation - openly published things are the same thing, minus the money.

The inventor's grace period under first to file changes still gives them/their university a year to file if they publish openly.
I know the frontier “labs” are holding back publications.

I don’t think it will last among researchers who think beyond production LLMs

Superficially it sounds like this could create a bit more of a move toward doing compaction on some continuous basis, or compacting in batches once you hit the context limit, rather than starting fresh with a summary and system prompt..

Feels like high fidelity, fast compaction could be a path to “solving” long context.

This looks promising. I've added it to my reading list.
This is big for long-horizon tasks
None of the compaction accuracies look impressive.
I think matching or exceeding the original cache at 20% compacted size is fairly impressive.
The original cache had 70% accuracy, and the alternatives were only worse.
It sounds like you looked at figure 1 but not figure 3.