In plain words: A Neural Turing Machine gives a neural network a separate memory it can read and write; a working implementation now solves three sequence tasks from the original paper. Starting that memory with small constant values trains about twice as fast as the next best starting values.
Abstract
Neural Turing Machines (NTMs) are an instance of Memory Augmented Neural Networks, a new class of recurrent neural networks which decouple computation from memory by introducing an external memory unit. NTMs have demonstrated superior performance over Long Short-Term Memory Cells in several sequence learning tasks. A number of open source implementations of NTMs exist but are unstable during training and/or fail to replicate the reported performance of NTMs. This paper presents the details of our successful implementation of a NTM. Our implementation learns to solve three sequential learning tasks from the original NTM paper. We find that the choice of memory contents initialization scheme is crucial in successfully implementing a NTM. Networks with memory contents initialized to small constant values converge on average 2 times faster than the next best memory contents initialization scheme.
Mark Collier, Joeran Beel
arXiv:1807.08518 · cs.LG, stat.ML · submitted Jul 23, 2018 · updated Jul 26, 2018
abstract · pdf · html
Also compared to the open source implementation (https://github.com/snowkylin/ntm) it seems like his main novel claim is that he looked at different memory initialisation patterns.
Edit:
compare the original: https://github.com/snowkylin/ntm/blob/master/ntm/ntm_cell.py
to the derivative work: https://github.com/MarkPKCollier/NeuralTuringMachine/blob/ma...
from what I can tell the main innovation is that the derivative work uses a named tuple instead of a dictionary for state keeping and there is new memory initialisation code. The original author apparently initialised the memory randomly. I also feel like the paper should cite the implementation they are basing their work on. The paper https://arxiv.org/pdf/1807.08518.pdf merely states that other implementations exist on page one and makes no mention of the fact that their implementation is based on one of those. Combine that with the fact that they are asking people in the Readme to cite their paper feels like not a very good idea.