In plain words: A sequence-reading network with built-in memory was wired directly onto a reprogrammable chip, so its step-by-step math runs in custom hardware. It ran over 21 times faster than the chip's own ARM processor core.
Abstract
Recurrent Neural Networks (RNNs) have the ability to retain memory and learn data sequences. Due to the recurrent nature of RNNs, it is sometimes hard to parallelize all its computations on conventional hardware. CPUs do not currently offer large parallelism, while GPUs offer limited parallelism due to sequential components of RNN models. In this paper we present a hardware implementation of Long-Short Term Memory (LSTM) recurrent network on the programmable logic Zynq 7020 FPGA from Xilinx. We implemented a RNN with $2$ layers and $128$ hidden units in hardware and it has been tested using a character level language model. The implementation is more than $21\times$ faster than the ARM CPU embedded on the Zynq 7020 FPGA. This work can potentially evolve to a RNN co-processor for future mobile devices.
Andre Xian Ming Chang, Berin Martini, Eugenio Culurciello
arXiv:1511.05552 · cs.NE · submitted Nov 17, 2015 · updated Mar 4, 2016
abstract · pdf · html · 7 pages, 8 figures, changed format, added figures, added references, modified introduction
I'm left curious on the performance gain factor when scaling the network in terms of layers and units. Would the performance gap widen as the RNN grows?