about
Cordic Is All You Need (arxiv.org)
2 points by PaulHoule on Apr 2, 2025 | hide | past | pdf | discuss on HN

In plain words: A chip design reuses one rotation-based math block to do both multiply-accumulate math and nonlinear steps like tanh and sigmoid, instead of separate circuits. It ran up to 4.64 times faster than earlier chips while cutting power and area, with a small accuracy drop.

Abstract · CORDIC Is All You Need

Artificial intelligence necessitates adaptable hardware accelerators for efficient high-throughput million operations. We present pipelined architecture with CORDIC block for linear MAC computations and nonlinear iterative Activation Functions (AF) such as $tanh$, $sigmoid$, and $softmax$. This approach focuses on a Reconfigurable Processing Engine (RPE) based systolic array, with 40\% pruning rate, enhanced throughput up to 4.64$\times$, and reduction in power and area by 5.02 $\times$ and 4.06 $\times$ at CMOS 28 nm, with minor accuracy loss. FPGA implementation achieves a reduction of up to 2.5 $\times$ resource savings and 3 $\times$ power compared to prior works. The Systolic CORDIC engine for Reconfigurability and Enhanced throughput (SYCore) deploys an output stationary dataflow with the CAESAR control engine for diverse AI workloads such as Transformers, RNNs/LSTMs, and DNNs for applications like image detection, LLMs, and speech recognition. The energy-efficient and flexible approach extends the enhanced approach for edge AI accelerators supporting emerging workloads.

Omkar Kokane, Adam Teman, Anushka Jha, Guru Prasath SL, Gopal Raut, Mukul Lokhande, S. V. Jaya Chand, Tanushree Dewangan, Santosh Kumar Vishvakarma
arXiv:2503.11685 · cs.AR, cs.CV, eess.IV · submitted Mar 4, 2025
abstract · pdf · html

add comment on HN