In plain words: Large memory networks for language keep a huge grid of numbers that is slow to train. Splitting that grid into two smaller pieces multiplied together, or into independent chunks, cuts the parameter count and trains much faster while reaching nearly the best word-prediction score.
Abstract
We present two simple ways of reducing the number of parameters and accelerating the training of large Long Short-Term Memory (LSTM) networks: the first one is "matrix factorization by design" of LSTM matrix into the product of two smaller matrices, and the second one is partitioning of LSTM matrix, its inputs and states into the independent groups. Both approaches allow us to train large LSTM networks significantly faster to the near state-of the art perplexity while using significantly less RNN parameters.
Oleksii Kuchaiev, Boris Ginsburg
arXiv:1703.10722 · cs.CL, cs.NE, stat.ML · submitted Mar 31, 2017 · updated Feb 24, 2018
abstract · pdf · html · accepted to ICLR 2017 Workshop
From my experience, RNN weights and the recurrent weights Tf in an LSTM tend to look more like (I + low_rank) rather than low_rank. To be more specific, I gather that with your F-LSTM you do:
Where W1 and W2 are "factorized by design".However, it seems like the recurrent weights (those for f in your paper) should look more like
That way, you are imposing the simplification that the Tf ~= (I + low_rank) instead of Tf ~= low_rank. Have you considered this?