In plain words: They slim down the LSTM memory cell by cutting some of the wiring—input, bias, and hidden signals—feeding its three gates. On two sequence tasks, these simpler versions matched the standard LSTM's accuracy while using fewer adjustable numbers.
Abstract · Simplified Gating in Long Short-term Memory (LSTM) Recurrent Neural Networks
The standard LSTM recurrent neural networks while very powerful in long-range dependency sequence applications have highly complex structure and relatively large (adaptive) parameters. In this work, we present empirical comparison between the standard LSTM recurrent neural network architecture and three new parameter-reduced variants obtained by eliminating combinations of the input signal, bias, and hidden unit signals from individual gating signals. The experiments on two sequence datasets show that the three new variants, called simply as LSTM1, LSTM2, and LSTM3, can achieve comparable performance to the standard LSTM model with less (adaptive) parameters.
Yuzhen Lu, Fathi M. Salem
arXiv:1701.03441 · cs.NE, stat.ML · submitted Jan 12, 2017
abstract · pdf · 5 pages, 4 Figures, 3 Tables. arXiv admin note: substantial text overlap with arXiv:1612.03707
I work in the field and I might read this later - but that's honestly only a might. The datasets they examine aren't particularly impactful or interesting and the paper is preliminary.
MNIST is a standard complained about dataset in vision (with someone recently noting it's more a unit test than a benchmark) but is infrequently used as a dataset for RNNs, other than potentially permuted MNIST which isn't used here. The IMDb dataset is at least standard for RNNs but also not representative of the complexity of many sequence tasks.
The primary statement being made is that the simpler LSTM1/2/3 can achieve similar numbers to that of the LSTM when using a proper hyper-parameter search. That's potentially useful to know but also likely not the thing stopping practitioners from putting such work into the field. Many architectures are also limited in the number of times they can be trained due to computational restrictions - otherwise I'd usually strongly suggest hyper-parameter search!
If people are interested in this type of analysis over RNN architectures, I recommend the older but still useful "An Empirical Exploration of Recurrent Network Architectures"[1]. The primary contribution there is that forget gates should be set to 1 for LSTMs, which was used and then forgotten for many years, but they do present various LSTM variants (MUT1/2/3) that are more computationally efficient too. These were integrated into Keras (a Python machine learning library) for some time. They also show their results over more datasets (arithmetic, XML modeling, language modeling on PTB, and music prediction) for a more convincing and nuanced discussion.
P.S. I'll note I saw this on Nando de Freitas' Twitter feed and assume that's why it was posted here (given he's a Professor of Computer Science at the University of Oxford and a senior researcher at Google). A retweet doesn't constitute an endorsement though, especially in science. I'm still confused as to why Hacker News, a very general crowd in tech, care particularly for one deep learning paper and not another. Color me confused :)
[1]: https://research.google.com/pubs/pub45473.html