In plain words: Instead of saving the attention memory each layer needs to remember earlier words, this stores it for only a few layers and reuses it, cutting memory use. It processed up to 26 times more text per second than standard transformers, with similar task performance.
Abstract · Layer-Condensed KV Cache for Efficient Inference of Large Language Models
Huge memory consumption has been a major bottleneck for deploying high-throughput large language models in real-world applications. In addition to the large number of parameters, the key-value (KV) cache for the attention mechanism in the transformer architecture consumes a significant amount of memory, especially when the number of layers is large for deep language models. In this paper, we propose a novel method that only computes and caches the KVs of a small number of layers, thus significantly saving memory consumption and improving inference throughput. Our experiments on large language models show that our method achieves up to 26$\times$ higher throughput than standard transformers and competitive performance in language modeling and downstream tasks. In addition, our method is orthogonal to existing transformer memory-saving techniques, so it is straightforward to integrate them with our model, achieving further improvement in inference efficiency. Our code is available at https://github.com/whyNLP/LCKV.
Haoyi Wu, Kewei Tu
arXiv:2405.10637 · cs.CL · submitted May 17, 2024 · updated Jun 4, 2024
abstract · pdf · html · Accepted to ACL2024 main conference
Initial result -- those KV caches in lower layers matter, and output suffered.
Updated plan -- cull half the KV layers! This works 'nearly' as well as keeping all of them, with memory and compute savings.
Downside - triple the training, worse out of band / long context performance.
This feels to me like a technique you'd use on a particular architecture deployed at the edge where compute matters and you have a little extra room on performance. Phi-3 on raspberry pi, basically.
Interesting! As always, I wish models showed prompt output in their papers, not just perplexity numbers. But, here we are.