about
Analog In-Memory Computing Attention Mechanism for Fast LLMs (arxiv.org)
4 points by bilsbie on Sep 12, 2025 | hide | past | pdf | discuss on HN

In plain words: A new chip stores each token's numbers in memory cells and computes attention right there, avoiding the GPU memory shuffle. A tuning trick avoids retraining and matches a well-known earlier language model's text quality, using up to 100,000 times less energy than GPUs.

Abstract · Analog In-Memory Computing Attention Mechanism for Fast and Energy-Efficient Large Language Models

Transformer networks, driven by self-attention, are central to Large Language Models. In generative Transformers, self-attention uses cache memory to store token projections, avoiding recomputation at each time step. However, GPU-stored projections must be loaded into SRAM for each new generation step, causing latency and energy bottlenecks. We present a custom self-attention in-memory computing architecture based on emerging charge-based memories called gain cells, which can be efficiently written to store new tokens during sequence generation and enable parallel analog dot-product computation required for self-attention. However, the analog gain cell circuits introduce non-idealities and constraints preventing the direct mapping of pre-trained models. To circumvent this problem, we design an initialization algorithm achieving text processing performance comparable to GPT-2 without training from scratch. Our architecture respectively reduces attention latency and energy consumption by up to two and five orders of magnitude compared to GPUs, marking a significant step toward ultra-fast, low-power generative Transformers.

Nathan Leroux, Paul-Philipp Manea, Chirag Sudarshan, Jan Finkbeiner, Sebastian Siegel, John Paul Strachan, Emre Neftci
arXiv:2409.19315 · cs.NE, cs.AI, cs.AR, cs.ET · submitted Sep 28, 2024 · updated Nov 25, 2024
abstract · pdf · html · 25 pages, 6 figures, 1 table

add comment on HN