In plain words: Sorbet swaps the transformer's two power-hungry steps—softmax and layer normalization—for simple bit-shifting versions, so a spiking language model can run on brain-like chips. It uses 27 times less energy than BERT while keeping similar accuracy.
Abstract · Sorbet: A Neuromorphic Hardware-Compatible Transformer-Based Spiking Language Model
For reasons such as privacy, there are use cases for language models at the edge. This has given rise to small language models targeted for deployment in resource-constrained devices where energy efficiency is critical. Spiking neural networks (SNNs) offer a promising solution due to their energy efficiency, and there are already works on realizing transformer-based models on SNNs. However, key operations like softmax and layer normalization (LN) are difficult to implement on neuromorphic hardware, and many of these early works sidestepped them. To address these challenges, we introduce Sorbet, a transformer-based spiking language model that is more neuromorphic hardware-compatible. Sorbet incorporates a novel shifting-based softmax called PTsoftmax and a Bit Shifting PowerNorm (BSPN), both designed to replace the respective energy-intensive operations. By leveraging knowledge distillation and model quantization, Sorbet achieved a highly compressed binary weight model that maintains competitive performance while achieving $27.16\times$ energy savings compared to BERT. We validate Sorbet through extensive testing on the GLUE benchmark and a series of ablation studies, demonstrating its potential as an energy-efficient solution for language model inference. Our code is publicly available at \href{https://github.com/Kaiwen-Tang/Sorbet}{https://github.com/Kaiwen-Tang/Sorbet}
Kaiwen Tang, Zhanglu Yan, Weng-Fai Wong
arXiv:2409.15298 · cs.NE, cs.CL, cs.LG · submitted Sep 4, 2024 · updated Jan 2, 2026
abstract · pdf · html · Accepted by ICML 2025. Camera-ready version
In particular the connection between the typical weighted-sum plus activation function and a simplistic spiking model where one considers the output simply by the spiking rate was illuminating (section 3).
[1]: https://www.ncbi.nlm.nih.gov/pmc/articles/PMC9313413/ Spiking Neural Networks and Their Applications: A Review