about
FPGA-based tiled matrix multiplication accelerator for self-attention (arxiv.org)
3 points by sha_rad 165 days ago | hide | past | pdf | discuss on HN

In plain words: A small reprogrammable chip runs the big matrix multiplications behind a language model's attention, keeping data on the chip and reusing it in blocks to avoid refetching. For DistilBERT's projections it ran 7 times faster than a phone-style processor, hitting 3.1 billion operations per second.

Abstract · Design and Implementation of an FPGA-Based Hardware Accelerator for Transformer

Transformer-based large language models (LLMs) rely heavily on intensive matrix multiplications for attention and feed-forward layers, with the Q, K, and V linear projections in the Multi-Head Self-Attention (MHA) module constituting a decisive performance bottleneck. In this work, we introduce a highly optimized tiled matrix multiplication accelerator on a resource-constrained Xilinx KV260 FPGA that not only addresses this challenge but sets a new standard for efficiency and performance. Our design exploits persistent on-chip storage, a robust two-level tiling strategy for maximal data reuse, and a systolic-like unrolled compute engine that together deliver unparalleled speed and energy efficiency. Integrated with DistilBERT for Q, K, and V projections, our accelerator achieves an unequivocal 7x speedup over ARM CPU implementations (PyTorch) and an extraordinary 200x improvement over naive NumPy, reaching a throughput of up to 3.1~GFLOPs for matrix multiplications on (64,768) x (768,3072) matrices while operating at a conservative 100 MHz. These results decisively demonstrate the transformative potential of FPGA-based acceleration for critical Transformer operations, paving the way for scalable and energy-efficient deep learning inference on edge devices.

Richie Li, Sicheng Chen
arXiv:2503.16731 · cs.AR, cs.CL, cs.LG · submitted Mar 20, 2025 · updated May 20, 2025
abstract · pdf · html · 7 pages, 4 figures, 2 tables. Prepared in ACM conference style. Preprint under review

add comment on HN