about
Exceeding Conventional FPGA Roofline Limit by LUT-Based Efficient Multiplication (arxiv.org)
1 point by PaulHoule on Dec 11, 2024 | hide | past | pdf | 1 comment on HN

In plain words: This design does multiplications with look-up tables—tiny memory blocks 100 times more plentiful than the chip's dedicated multipliers—to speed up inference on FPGAs. It recognized images faster than any other FPGA accelerator, at 1627 per second with 70.95% top-1 accuracy on ImageNet.

Abstract · LUTMUL: Exceed Conventional FPGA Roofline Limit by LUT-based Efficient Multiplication for Neural Network Inference

For FPGA-based neural network accelerators, digital signal processing (DSP) blocks have traditionally been the cornerstone for handling multiplications. This paper introduces LUTMUL, which harnesses the potential of look-up tables (LUTs) for performing multiplications. The availability of LUTs typically outnumbers that of DSPs by a factor of 100, offering a significant computational advantage. By exploiting this advantage of LUTs, our method demonstrates a potential boost in the performance of FPGA-based neural network accelerators with a reconfigurable dataflow architecture. Our approach challenges the conventional peak performance on DSP-based accelerators and sets a new benchmark for efficient neural network inference on FPGAs. Experimental results demonstrate that our design achieves the best inference speed among all FPGA-based accelerators, achieving a throughput of 1627 images per second and maintaining a top-1 accuracy of 70.95% on the ImageNet dataset.

Yanyue Xie, Zhengang Li, Dana Diaconu, Suranga Handagala, Miriam Leeser, Xue Lin
arXiv:2411.11852 · cs.AR, cs.AI, cs.LG · submitted Nov 1, 2024
abstract · pdf · html · Accepted by ASPDAC 2025

add comment on HN

> ASPDAC ’25, January 20–23, 2025, Tokyo, Japan

> Experimental results demonstrate that our design maintains a top-1 accuracy of 70.95% on the ImageNet dataset and achieves a throughput of 1627 images per second on a single Alveo U280 FPGA, outperforming other FPGA-based MobileNet accelerators.

Well, other accelerators are smaller with fewer resources and apparently cheaper.

The conference is yet to be held, but the hardware is already obsolete.

Alveo U280 (launched 2019) is obsolete according to digikey and not mentioned on [1], page [2] is gone.

V100 GPU is outdated (launched 2017) either. Last prices (with VAT):

- V100: 7000 / 8900 / 9800 Euro (seems to be still available)

- U280: 10200 Euro (only 16 GB instead of 32 GB) (seems to be not easy available)

vs. prices from the article:

- V100: $11458

- U280: $7717

Looking at performance:

- nVidia V100 GPU: 112TFLOPs (FP16 Tensor)

- AMD RX 7900 XTX GPU: 131 TFLOPs (FP16) (not obsolete!) for 1000 Euro

Comparing an FPGA solution to an affordable and very easy available GPU would be more practical.

[1] https://www.amd.com/en/products/accelerators/alveo.html

[2] https://www.xilinx.com/products/boards-and-kits/alveo/u280.h...