In plain words: It converts models that store memory per attention group into DeepSeek's compressed attention style, letting DeepSeek's fast inference code run them. On LLaMA-2-7B it shrank the memory cache and ran 10.6x faster than the original at long context, with output quality back after fine-tuning.
Abstract · TransMLA: Multi-Head Latent Attention Is All You Need
In this paper, we present TransMLA, a framework that seamlessly converts any GQA-based pre-trained model into an MLA-based model. Our approach enables direct compatibility with DeepSeek's codebase, allowing these models to fully leverage DeepSeek-specific optimizations such as vLLM and SGlang. By compressing 93% of the KV cache in LLaMA-2-7B, TransMLA achieves a 10.6x inference speedup at an 8K context length while preserving meaningful output quality. Additionally, the model requires only 6 billion tokens for fine-tuning to regain performance on par with the original across multiple benchmarks. TransMLA offers a practical solution for migrating GQA-based models to the MLA structure. When combined with DeepSeek's advanced features, such as FP8 quantization and Multi-Token Prediction, even greater inference acceleration can be realized.
Fanxu Meng, Pingzhi Tang, Xiaojuan Tang, Zengwei Yao, Xing Sun, Muhan Zhang
arXiv:2502.07864 · cs.LG, cs.AI · submitted Feb 11, 2025 · updated Jun 12, 2025
abstract · pdf · html · https://github.com/fxmeng/TransMLA
[3.3] For saving the KV cache, only the intermediate latent representations need to be stored: [latex] where r is much smaller than nh · dh [n-sub-h, d-sub-h]
[background] In traditional multi-head attention you must cache full key and value matrices of size T x (nh · dh) where T is the token length, nh is the number of attention heads, dh is the dimensionality of each individual head
sounds like a big win for memory constrained environments like local inference