In plain words: A translation system that converts a sentence by passing it through stacked memory layers, reading and writing intermediate versions at each step. The deeper stack beat the best neural translator then available and nearly matched a traditional phrase-based system with a small word list.
Abstract · A Deep Memory-based Architecture for Sequence-to-Sequence Learning
We propose DEEPMEMORY, a novel deep architecture for sequence-to-sequence learning, which performs the task through a series of nonlinear transformations from the representation of the input sequence (e.g., a Chinese sentence) to the final output sequence (e.g., translation to English). Inspired by the recently proposed Neural Turing Machine (Graves et al., 2014), we store the intermediate representations in stacked layers of memories, and use read-write operations on the memories to realize the nonlinear transformations between the representations. The types of transformations are designed in advance but the parameters are learned from data. Through layer-by-layer transformations, DEEPMEMORY can model complicated relations between sequences necessary for applications such as machine translation between distant languages. The architecture can be trained with normal back-propagation on sequenceto-sequence data, and the learning can be easily scaled up to a large corpus. DEEPMEMORY is broad enough to subsume the state-of-the-art neural translation model in (Bahdanau et al., 2015) as its special case, while significantly improving upon the model with its deeper architecture. Remarkably, DEEPMEMORY, being purely neural network-based, can achieve performance comparable to the traditional phrase-based machine translation system Moses with a small vocabulary and a modest parameter size.
Fandong Meng, Zhengdong Lu, Zhaopeng Tu, Hang Li, Qun Liu
arXiv:1506.06442 · cs.CL, cs.LG, cs.NE · submitted Jun 22, 2015 · updated Jan 7, 2016
abstract · pdf · html · 13 pages, Under review as a conference paper at ICLR 2016
http://papers.nips.cc/paper/5346-sequence-to-sequence-learni...
which "uses a multilayered Long Short-Term Memory (LSTM) to map the input sequence to a vector of a fixed dimensionality, and then another deep LSTM to decode the target sequence from the vector." The task was English to French.
Meng et al.(2015)(OP) translate a Chinese sequence to English, using a network based on Neural Turing Machines(NTM) which uses LTSM units, they name this novel architecture Neural Transformation Machine (NTRam).
The Neural Turing Machine(NTM) was proposed by Deepmind's Alex Graves, Greg Wayne & Ivo Danihelka, it couples a neural net and LTSM memory to produce a differentiable, thus trainable analogy to a Turing Machine or Von Neumann architecture - to perfom copying, sorting and associative recall. An exploration of whether Neural Networks can be put to basic computing functions.
Neural Turing Machines(2014) by Alex Graves, Greg Wayne, Ivo Danihelka
http://arxiv.org/abs/1410.5401v2