In plain words: Instead of pruning RNN weights into a sparse pattern, this stores weight blocks as a repeating circular pattern, which shrinks the model and speeds up computation. On an FPGA it was up to 35.7 times more energy-efficient than pruning, with barely any accuracy loss.
Abstract
Recurrent Neural Networks (RNNs) are becoming increasingly important for time series-related applications which require efficient and real-time implementations. The recent pruning based work ESE suffers from degradation of performance/energy efficiency due to the irregular network structure after pruning. We propose block-circulant matrices for weight matrix representation in RNNs, thereby achieving simultaneous model compression and acceleration. We aim to implement RNNs in FPGA with highest performance and energy efficiency, with certain accuracy requirement (negligible accuracy degradation). Experimental results on actual FPGA deployments shows that the proposed framework achieves a maximum energy efficiency improvement of 35.7$\times$ compared with ESE.
Zhe Li, Shuo Wang, Caiwen Ding, Qinru Qiu, Yanzhi Wang, Yun Liang
arXiv:1803.07661 · cs.LG, math.NA, stat.ML · submitted Mar 20, 2018 · updated Mar 22, 2018
abstract · pdf · html · To appear in International Conference on Learning Representations 2018 Workshop Track
This reduces W's memory consumption by a factor of n, and makes other gains in computational efficiency possible. Read the paper for details.
However, as far as I can tell, it appears the authors have tested this technique only with one RNN architecture, in one task. It's hard to know whether the technique will hold up well in a broad range of RNN architectures/tasks.
[a] https://en.wikipedia.org/wiki/Circulant_matrix