In plain words: A framework describing how a memory network's units connect over time, plus three numbers scoring loop depth, step depth, and how fast information skips ahead. Tests suggest deeper loops and steps help, and letting information skip ahead helps most on tasks needing long memory.
Abstract
In this paper, we systematically analyze the connecting architectures of recurrent neural networks (RNNs). Our main contribution is twofold: first, we present a rigorous graph-theoretic framework describing the connecting architectures of RNNs in general. Second, we propose three architecture complexity measures of RNNs: (a) the recurrent depth, which captures the RNN's over-time nonlinear complexity, (b) the feedforward depth, which captures the local input-output nonlinearity (similar to the "depth" in feedforward neural networks (FNNs)), and (c) the recurrent skip coefficient which captures how rapidly the information propagates over time. We rigorously prove each measure's existence and computability. Our experimental results show that RNNs might benefit from larger recurrent depth and feedforward depth. We further demonstrate that increasing recurrent skip coefficient offers performance boosts on long term dependency problems.
Saizheng Zhang, Yuhuai Wu, Tong Che, Zhouhan Lin, Roland Memisevic, Ruslan Salakhutdinov, Yoshua Bengio
arXiv:1602.08210 · cs.LG, cs.NE · submitted Feb 26, 2016 · updated Nov 12, 2016
abstract · pdf · html · 17 pages, 8 figures; To appear in NIPS2016