about
Challenges and Research Directions for Large Language Model Inference Hardware (arxiv.org)
2 points by matt_d 267 days ago | hide | past | pdf | discuss on HN

In plain words: Generating text one word at a time makes chips run out of memory and move data between chips instead of doing math. It proposes four ideas, like wide-path flash storage and memory stacked on logic, to hold 10 times more data at near-fastest memory speeds.

Abstract

Large Language Model (LLM) inference is hard. The autoregressive Decode phase of the underlying Transformer model makes LLM inference fundamentally different from training. Exacerbated by recent AI trends, the primary challenges are memory and interconnect rather than compute. To address these challenges, we highlight four architecture research opportunities: High Bandwidth Flash for 10X memory capacity with HBM-like bandwidth; Processing-Near-Memory and 3D memory-logic stacking for high memory bandwidth; and low-latency interconnect to speedup communication. While our focus is datacenter AI, we also review their applicability for mobile devices.

Xiaoyu Ma, David Patterson
arXiv:2601.05047 · cs.AR, cs.AI, cs.LG · submitted Jan 8, 2026 · updated Feb 6, 2026
abstract · pdf · Accepted for publication by IEEE Computer, 2026

add comment on HN
Also discussed: Jan 2026 (123 points, 23 comments)