about
Marca: Mamba Accelerator with ReConfigurable Architecture (arxiv.org)
2 points by PaulHoule on Sep 23, 2024 | hide | past | pdf | discuss on HN

In plain words: A chip design for Mamba AI models reuses the same computing blocks for matrix math, simple per-value math, and nonlinear functions, while buffers keep shared data close by. It ran up to 11.66 times faster than a high-end GPU while using far less energy.

Abstract · MARCA: Mamba Accelerator with ReConfigurable Architecture

We propose a Mamba accelerator with reconfigurable architecture, MARCA.We propose three novel approaches in this paper. (1) Reduction alternative PE array architecture for both linear and element-wise operations. For linear operations, the reduction tree connected to PE arrays is enabled and executes the reduction operation. For element-wise operations, the reduction tree is disabled and the output bypasses. (2) Reusable nonlinear function unit based on the reconfigurable PE. We decompose the exponential function into element-wise operations and a shift operation by a fast biased exponential algorithm, and the activation function (SiLU) into a range detection and element-wise operations by a piecewise approximation algorithm. Thus, the reconfigurable PEs are reused to execute nonlinear functions with negligible accuracy loss.(3) Intra-operation and inter-operation buffer management strategy. We propose intra-operation buffer management strategy to maximize input data sharing for linear operations within operations, and inter-operation strategy for element-wise operations between operations. We conduct extensive experiments on Mamba model families with different sizes.MARCA achieves up to 463.22$\times$/11.66$\times$ speedup and up to 9761.42$\times$/242.52$\times$ energy efficiency compared to Intel Xeon 8358P CPU and NVIDIA Tesla A100 GPU implementations, respectively.

Jinhao Li, Shan Huang, Jiaming Xu, Jun Liu, Li Ding, Ningyi Xu, Guohao Dai
arXiv:2409.11440 · cs.AR, cs.AI · submitted Sep 16, 2024
abstract · pdf · html · 9 pages, 10 figures, accepted by ICCAD 2024. arXiv admin note: text overlap with arXiv:2001.02514 by other authors

add comment on HN