about
SpeedLLM: An FPGA Co-Design of Large Language Model Inference Accelerator (arxiv.org)
2 points by juanviera23 on Aug 1, 2025 | hide | past | pdf | discuss on HN

In plain words: A custom chip design for running a small language model on edge devices streams data through a tight read-compute-write pipeline, reuses on-chip memory, and merges operations to cut wasted work. It ran up to 4.8 times faster than the standard Tinyllama setup while using 1.18 times less energy.

Abstract · SpeedLLM: An FPGA Co-design of Large Language Model Inference Accelerator

This paper introduces SpeedLLM, a neural network accelerator designed on the Xilinx Alevo U280 platform and optimized for the Tinyllama framework to enhance edge computing performance. Key innovations include data stream parallelism, a memory reuse strategy, and Llama2 operator fusion, which collectively reduce latency and energy consumption. SpeedLLM's data pipeline architecture optimizes the read-compute-write cycle, while the memory strategy minimizes FPGA resource demands. The operator fusion boosts computational density and throughput. Results show SpeedLLM outperforms traditional Tinyllama implementations, achieving up to 4.8* faster performance and 1.18* lower energy consumption, offering improvements in edge devices.

Peipei Wang, Wu Guan, Liping Liang, Zhijun Wang, Hanqing Luo, Zhibin Zhang
arXiv:2507.14139 · cs.AR · submitted May 7, 2025
abstract · pdf · html

add comment on HN