In plain words: A language model that skips the matrix multiplications of usual models runs on Intel's brain-inspired Loihi 2 chip, which calculates only when signals change. Tests show it runs up to three times faster and uses less energy than transformer models on a small GPU.
Abstract
Large language models (LLMs) deliver impressive performance but require large amounts of energy. In this work, we present a MatMul-free LLM architecture adapted for Intel's neuromorphic processor, Loihi 2. Our approach leverages Loihi 2's support for low-precision, event-driven computation and stateful processing. Our hardware-aware quantized model on GPU demonstrates that a 370M parameter MatMul-free model can be quantized with no accuracy loss. Based on preliminary results, we report up to 3x higher throughput with 2x less energy, compared to transformer-based LLMs on an edge GPU, with significantly better scaling. Further hardware optimizations will increase throughput and decrease energy consumption. These results show the potential of neuromorphic hardware for efficient inference and pave the way for efficient reasoning models capable of generating complex, long-form text rapidly and cost-effectively.
Steven Abreu, Sumit Bam Shrestha, Rui-Jie Zhu, Jason Eshraghian
arXiv:2503.18002 · cs.NE, cs.AI, cs.AR, cs.LG · submitted Feb 12, 2025 · updated Mar 25, 2025
abstract · pdf · html · Accepted to International Conference on Learning Representations (ICLR) Workshop on Scalable Optimization for Efficient and Adaptive Foundation Models (SCOPE)