In plain words: Generating text one word at a time makes chips run out of memory and move data between chips instead of doing math. It proposes four ideas, like wide-path flash storage and memory stacked on logic, to hold 10 times more data at near-fastest memory speeds.
Abstract
Large Language Model (LLM) inference is hard. The autoregressive Decode phase of the underlying Transformer model makes LLM inference fundamentally different from training. Exacerbated by recent AI trends, the primary challenges are memory and interconnect rather than compute. To address these challenges, we highlight four architecture research opportunities: High Bandwidth Flash for 10X memory capacity with HBM-like bandwidth; Processing-Near-Memory and 3D memory-logic stacking for high memory bandwidth; and low-latency interconnect to speedup communication. While our focus is datacenter AI, we also review their applicability for mobile devices.
Xiaoyu Ma, David Patterson
arXiv:2601.05047 · cs.AR, cs.AI, cs.LG · submitted Jan 8, 2026 · updated Feb 6, 2026
abstract · pdf · Accepted for publication by IEEE Computer, 2026