about
Simplified Gating in Long Short-Term Memory (LSTM) Recurrent Neural Networks (arxiv.org)
49 points by MichaelBurge on Jan 13, 2017 | hide | past | pdf | 11 comments on HN

In plain words: They slim down the LSTM memory cell by cutting some of the wiring—input, bias, and hidden signals—feeding its three gates. On two sequence tasks, these simpler versions matched the standard LSTM's accuracy while using fewer adjustable numbers.

Abstract · Simplified Gating in Long Short-term Memory (LSTM) Recurrent Neural Networks

The standard LSTM recurrent neural networks while very powerful in long-range dependency sequence applications have highly complex structure and relatively large (adaptive) parameters. In this work, we present empirical comparison between the standard LSTM recurrent neural network architecture and three new parameter-reduced variants obtained by eliminating combinations of the input signal, bias, and hidden unit signals from individual gating signals. The experiments on two sequence datasets show that the three new variants, called simply as LSTM1, LSTM2, and LSTM3, can achieve comparable performance to the standard LSTM model with less (adaptive) parameters.

Yuzhen Lu, Fathi M. Salem
arXiv:1701.03441 · cs.NE, stat.ML · submitted Jan 12, 2017
abstract · pdf · 5 pages, 4 Figures, 3 Tables. arXiv admin note: substantial text overlap with arXiv:1612.03707

add comment on HN

To the Hacker News community, as I'm genuinely curious, does this link contribute anything to you or your understanding of deep learning? 29 people voted for it, zero comments, and it's still on the front page after five hours. It's a very technical and very specific paper that has not yet had any real analysis by the broader ML community. From the previous discussions on deep learning, I generally suspect that seeing "neural networks" is an upvote trigger but then very little good discussion continues past that.

I work in the field and I might read this later - but that's honestly only a might. The datasets they examine aren't particularly impactful or interesting and the paper is preliminary.

MNIST is a standard complained about dataset in vision (with someone recently noting it's more a unit test than a benchmark) but is infrequently used as a dataset for RNNs, other than potentially permuted MNIST which isn't used here. The IMDb dataset is at least standard for RNNs but also not representative of the complexity of many sequence tasks.

The primary statement being made is that the simpler LSTM1/2/3 can achieve similar numbers to that of the LSTM when using a proper hyper-parameter search. That's potentially useful to know but also likely not the thing stopping practitioners from putting such work into the field. Many architectures are also limited in the number of times they can be trained due to computational restrictions - otherwise I'd usually strongly suggest hyper-parameter search!

If people are interested in this type of analysis over RNN architectures, I recommend the older but still useful "An Empirical Exploration of Recurrent Network Architectures"[1]. The primary contribution there is that forget gates should be set to 1 for LSTMs, which was used and then forgotten for many years, but they do present various LSTM variants (MUT1/2/3) that are more computationally efficient too. These were integrated into Keras (a Python machine learning library) for some time. They also show their results over more datasets (arithmetic, XML modeling, language modeling on PTB, and music prediction) for a more convincing and nuanced discussion.

P.S. I'll note I saw this on Nando de Freitas' Twitter feed and assume that's why it was posted here (given he's a Professor of Computer Science at the University of Oxford and a senior researcher at Google). A retweet doesn't constitute an endorsement though, especially in science. I'm still confused as to why Hacker News, a very general crowd in tech, care particularly for one deep learning paper and not another. Color me confused :)

[1]: https://research.google.com/pubs/pub45473.html

This should not be on HN for two reasons:

1) Only people how already know and used RNN could be interested

2) It is not a good paper

Let me explain. MNIST is not a good dataset for testing LSTMs. LSTMs were designed to handle sequential data. Of course you can use them with MNIST, but simple feed-forward network would be much easier and faster to train. RNNs give no advantage here. IMDb dataset is a decent choice for tests, but there is no convincing argument (or result) that makes LSTM1/LSTM2/LSTM3 a good choice. "Best accuracies of different LSTMs" table tells us absolutely nothing! You can just cherry-pick best results ignoring others. Other thing is authors didn't test GRU. GRUs are based on LSTM design, but simplified a bit. GRUs are known to perform better in many cases and are faster to train.

The paper recommended by user Smerity is the best paper comparing RNN cell architectures and their performance I know and I recommend it too ("An Empirical Exploration of Recurrent Network Architectures")

A better dataset for testing RNNs in general and LSTMs in particular would be the UCR Time Series Classification Archive [0].

[0] http://www.cs.ucr.edu/~eamonn/time_series_data

I would also mention that the method used to design the new LSTM units is much less interesting than another recent paper on designing LSTM units, using reinforcement learning NNs: Zoph & Le 2016 http://openreview.net/pdf?id=r1Ue8Hcxg "Neural architecture search with reinforcement learning" (in addition, tested on a more relevant sequence task and performs better, not just similarly).
To be fair, if it hadn't been upvoted we wouldn't have got your awesome response.
First, nawwwww, thanks :)

Second though, I've sadly stopped responding (or even reading) many of the machine learning and especially broader artificial intelligence posts on Hacker News. The discussions are rarely discussions that one can contribute to constructively meaning the time spent rarely provides a positive return. My primary social channel for ML/AI is Twitter - a great community but a depressingly inadequate tool for such discussions. For many other topics however Hacker News is still my goto community.

It's a shame as I love to discuss machine learning and artificial intelligence but, amongst other things, there's a Godwin style law that these discussions inevitably lead to comparisons to the human brain and/or the singularity and/or killer robots and/or "no bias" in machine learning.

Yes, people on HN are very enthusiastic about ANNs and Deep Learning, because they've heard it's a big technological revolution and they want to keep informed. In their enthusiasm, some say things that don't make sense (my favourite is a comment that proposed to train an ANN to acquire an imagination- presumably by training it on pictures of unicorns and elves as examples thereof).

On the other hand, I think there's also enough people on HN who know enough about those subjects, that the comments that get the most upvotes (and, consequently, rise to the top) are the ones that make the most sense.

Edit: to be fair, all the stuff about brains and singularities is not the fault of HN users. Very prominent researchers (Hinton, Schmidhuber, others) keep pushing the "based on the rain" line, for instance.

What do you think about r/machinelearning? That community is much more informed, but there are still very few comments probably because it's much smaller than HN.
You don't have to respond to comments that don't interest you.
I'm the submitter. I usually find it helpful to vary the hyperparameters to see how they change a network's performance, and hadn't yet done that for LSTMs. This one had a few ideas for where to start, so I saved it for later. Your paper looks like a much better reference, though.

I'm a little surprised it got so many votes, too. I would've liked to see discussion on the Spatial Transformer Networks paper I posted yesterday, which seemed much more useful.

I read and generally upvote any novel work on ANNs regardless of practical use, because I feel like the ML field is still just as much art as science and inspiration begets creativity. At least for me, it might take weeks of experimentation before I have anything constructive to add to the discussion.