In plain words: Language models usually rely on matrix multiplication in their attention and feed-forward layers; this design swaps those heavy multiplications for simpler per-number operations. It matches standard pre-trained Transformers at billion-parameter scale while cutting training memory by up to 61%.
Abstract · Scalable MatMul-free Language Modeling
Large Language Models (LLMs) have fundamentally altered how we approach scaling in machine learning. However, these models pose substantial computational and memory challenges, primarily due to the reliance on matrix multiplication (MatMul) within their attention and feed-forward (FFN) layers. We demonstrate that MatMul operations can be eliminated from LLMs while maintaining strong performance, even at billion-parameter scales. Our MatMul-free models, tested on models up to 2.7B parameters, are comparable to state-of-the-art pre-trained Transformers, and the performance gap narrows as model size increases. Our approach yields significant memory savings: a GPU-efficient implementation reduces memory consumption by up to 61% during training and over 10x during inference. When adapted for a multi-chip neuromorphic system, the model leverages asynchronous processing to achieve 4x higher throughput with 10x less energy than edge GPUs.
Rui-Jie Zhu, Yu Zhang, Steven Abreu, Ethan Sifferman, Tyler Sheaves, Yiqiao Wang, Dustin Richmond, Sumit Bam Shrestha, Peng Zhou, Jason K. Eshraghian
arXiv:2406.02528 · cs.CL · submitted Jun 4, 2024 · updated Jul 25, 2025
abstract · pdf · html
https://arxiv.org/abs/2402.17764
The main addition of the new paper seems to be the implementation of optimized and fused kernels using triton, as seen here:
https://github.com/ridgerchu/matmulfreellm/blob/master/mmfre...
This is quite useful, as this should make training this type of LLMs much more efficient.
So this is a ternary weight LLM using quantization aware training (QAT). The activations are quantized to 8 bits. The matmal is still there, but it is multiplying the 8 bit activations by one bit values.
Quantization aware training with low bit weights seems to lead to reduced overfitting by an intrensic tendency to regularize. However, also the model capacity should be reduced compared to a model with the same number of weights and a higher number of bits per weights. It's quite possible that this only becomes apparent after the models have been trained with a significant number of tokens, as LLMs seem to be quite sparse.
Edit: In addition to the QAT they also changed the model architecture to use a linear transformer to reduce reliance on multiplications in the attention mechanism. Thanks to logicchains for pointing this out.