about
Factorization tricks for LSTM networks (arxiv.org)
69 points by Katydid on Apr 10, 2017 | hide | past | pdf | 14 comments on HN

In plain words: Large memory networks for language keep a huge grid of numbers that is slow to train. Splitting that grid into two smaller pieces multiplied together, or into independent chunks, cuts the parameter count and trains much faster while reaching nearly the best word-prediction score.

Abstract

We present two simple ways of reducing the number of parameters and accelerating the training of large Long Short-Term Memory (LSTM) networks: the first one is "matrix factorization by design" of LSTM matrix into the product of two smaller matrices, and the second one is partitioning of LSTM matrix, its inputs and states into the independent groups. Both approaches allow us to train large LSTM networks significantly faster to the near state-of the art perplexity while using significantly less RNN parameters.

Oleksii Kuchaiev, Boris Ginsburg
arXiv:1703.10722 · cs.CL, cs.NE, stat.ML · submitted Mar 31, 2017 · updated Feb 24, 2018
abstract · pdf · html · accepted to ICLR 2017 Workshop

add comment on HN

Since the authors appear to be listening on here, I have a question about the method.

From my experience, RNN weights and the recurrent weights Tf in an LSTM tend to look more like (I + low_rank) rather than low_rank. To be more specific, I gather that with your F-LSTM you do:

  T1 = W1 * input
  T2 = W2 * T1
  output = T2
  so
  output = W2 * W1 * input
Where W1 and W2 are "factorized by design".

However, it seems like the recurrent weights (those for f in your paper) should look more like

  T1 = W1 * input
  T2 = W2 * T1
  output = input + T2
  so
  output = (I + W2 * W1) * input
That way, you are imposing the simplification that the Tf ~= (I + low_rank) instead of Tf ~= low_rank. Have you considered this?
>>"RNN weights and the recurrent weights Tf in an LSTM tend to look more like (I + low_rank)" - Do you have reference for this? But thanks for suggestion, I'll have a look into this.

>>"I gather that with your F-LSTM you do:" - Your understanding of F-LSTM looks correct

>>"output = input + T2" Not quite clear where to put non-linearities. But this looks similar to residual connections. Which, in my experience, is almost always a good idea.

If you've seen tensor trains, might be interesting to reduce the rank even further by splitting an input into many sequential small matrices (instead of just two.)
It's crazy and scary at the same time how fast new approaches are improving machine learning efficiency. Am I only one who thinks that we will be able to simulate the brain with much less computational power than the brain itself has?
The brain is a highly inefficient system evolved arbitrarily and slowly over a long time. It does seem likely that we could, with much fewer computational resources, simulate the function of some of the tasks which brains are responsible for or "good at", such as image classification.

However, "simulating the brain" in general is a much broader problem in scope. Consider: Can you even really truly say that you have "simulated the brain" without also including a physical body for this system, since a brain is inherently tied to a body and corporal existence in the world? It's not clear.

It's likely the problem won't be computational power but figuring out how to extract or isolate only the bits of the system we are interested in from the rest. Consider the genome, for instance: we have had the human genome mapped since 2003 but the unfathomable complexity of the system as a whole makes it difficult (though not impossible) to apply for useful stuff.

It's not just about computational power, it's also about defining computational models that can create human-competitive performance for certain tasks, or, in the case of AGI, for all tasks that might be expected of a human. That's the fiendishly tricky bit.

Who you calling inefficient? :-p

Those "inefficiencies" are for handling a messy noisy world where we can't just settle on a single solution. There is allot of overhead to play in the game of evolution.

author here. didn't expect our paper to appear on HN :) But regarding: "Am I only one who thinks that we will be able to simulate the brain with much less computational power than the brain itself has?" - I really doubt that. In fact my bet is that we won't. If anything our learning approaches looks much more inefficient (data-wise) compared to what brain can do. - And the paper doesn't study this question. It is a focused study on speed/accuracy tradeoff for LSTM cells
You're not the only one.

A neuron nucleus can be as small as 3 μm, ... which could pack about tens or hundreds million 5nm transistors.

There is plenty of room at the bottom, as Feynman said.

What does "computation power of the brain mean"? The brain does not operate similar to a computer; how can you make comparisons?
I think communication bandwidth is the best estimator of the brain capacity.
How so?
Seems quite interesting. I would try it this week.
author here. let me know (on Github) if you encounter any issues.
A much better paper worthy of attention from the HN crowd is this: Using Human Brain Activity to Guide Machine Learning (https://arxiv.org/abs/1703.05463)