about
Efficient Recurrent Neural Networks using Structured Matrices in FPGAs (arxiv.org)
79 points by godelmachine on Mar 22, 2018 | hide | past | pdf | 19 comments on HN

In plain words: Instead of pruning RNN weights into a sparse pattern, this stores weight blocks as a repeating circular pattern, which shrinks the model and speeds up computation. On an FPGA it was up to 35.7 times more energy-efficient than pruning, with barely any accuracy loss.

Abstract

Recurrent Neural Networks (RNNs) are becoming increasingly important for time series-related applications which require efficient and real-time implementations. The recent pruning based work ESE suffers from degradation of performance/energy efficiency due to the irregular network structure after pruning. We propose block-circulant matrices for weight matrix representation in RNNs, thereby achieving simultaneous model compression and acceleration. We aim to implement RNNs in FPGA with highest performance and energy efficiency, with certain accuracy requirement (negligible accuracy degradation). Experimental results on actual FPGA deployments shows that the proposed framework achieves a maximum energy efficiency improvement of 35.7$\times$ compared with ESE.

Zhe Li, Shuo Wang, Caiwen Ding, Qinru Qiu, Yanzhi Wang, Yun Liang
arXiv:1803.07661 · cs.LG, math.NA, stat.ML · submitted Mar 20, 2018 · updated Mar 22, 2018
abstract · pdf · html · To appear in International Conference on Learning Representations 2018 Workshop Track

add comment on HN

The main idea is to take each weight matrix W, divide it into smaller blocks, each of dimension, say, n×n, and make each block a circulant matrix that can be specified by a vector of only n elements.[a]

This reduces W's memory consumption by a factor of n, and makes other gains in computational efficiency possible. Read the paper for details.

However, as far as I can tell, it appears the authors have tested this technique only with one RNN architecture, in one task. It's hard to know whether the technique will hold up well in a broad range of RNN architectures/tasks.

[a] https://en.wikipedia.org/wiki/Circulant_matrix

Speech recognition be the task they implemented this for.
What about image recognition? Too complex? Or doable at 320x240 resolution?
How does this compare to GPU? These FPGAs are high end ones and not cheap for sure.
According to [1], a high-end FPGA can do "AlexNet inference performance: int16 over 2,400 img/s, int8 over 4,500 img/s.". Plug in your favorite nVidia numbers and compare.

[1] http://mipsology.com/zebra.html

Thank you very much for the link!
Hi Insru - how do you think GPUs would fare taking into consideration alain94040 's metrics?
Please look at this pdf: https://www.nvidia.com/content/tegra/embedded-systems/pdf/je... There is a table on page 9 and it claims, that Titan X can make 3216 img/s while TUL-KU115 FPGA board can make only >1000 img/s. This KU115 chip has a project price of ~$2000. Real time capability is nice, but that’s still expensive.
> This KU115 chip has a project price of ~$2000.

That seems odd - just the XCKU115-2FLVB2104E FPGA goes for $9000 on DigiKey.

Thanks for this PDF :)

Do you think TitanX has real time capability?

I am FPGA developer and have not much knowledge about GPUs. I just can make a guess, that no one can guarantee execution time on GPU. There are all little details well defined in FPGA design and result is always available after defined delay.
Though wouldn't it be far more interesting to compare the latencies at these rates, than the throughput? Roughly similar price and throughput, but very different energy efficiency and possibly latency could be rather significant for quite a few applications.
Xilinx has prepared some marketing material regarding latency: https://forums.xilinx.com/t5/Xcell-Daily-Blog/Xilinx-reVISIO...
For some reasons the Xilinx link is not opening.
Got it! Thanks :)
Regarding GPU, they have only said - "The ESE achieves higher energy efficiency than GPU, but its performance is lower and it cannot operate in real time"

So maybe we can assume the performance is quite fast - as compared to the GPU they used in ESE, and this work can operate in real time?

Table 2 has latencys (measured in microseconds) and a frames-per-second measure too (which places it a ~250k-400k fps). According to them, ESE has a 17k fps "theoretical max", but in reality has a lower performance.

I don't really get the fps measure -- without context, I can't say above which threshold they'd be able to perform real-time operation (this is surely due to my lack of knowledge in this domain).

From the original paper: ESE outperforms Core i7 CPU and Pascal Titan X GPU by factors of 43× and 3× on speed, and it is 40× and 11.5× more energy efficient than the CPU and GPU respectively.