In plain words: A network that reads from and writes to an external memory turns written words into how they should be spoken, without being tied to one language. It makes fewer mistakes than usual step-by-step recurrent systems, while needing less data, training time, and processing power.
Abstract · Text normalization using memory augmented neural networks
We perform text normalization, i.e. the transformation of words from the written to the spoken form, using a memory augmented neural network. With the addition of dynamic memory access and storage mechanism, we present a neural architecture that will serve as a language-agnostic text normalization system while avoiding the kind of unacceptable errors made by the LSTM-based recurrent neural networks. By successfully reducing the frequency of such mistakes, we show that this novel architecture is indeed a better alternative. Our proposed system requires significantly lesser amounts of data, training time and compute resources. Additionally, we perform data up-sampling, circumventing the data sparsity problem in some semiotic classes, to show that sufficient examples in any particular class can improve the performance of our text normalization system. Although a few occurrences of these errors still remain in certain semiotic classes, we demonstrate that memory augmented networks with meta-learning capabilities can open many doors to a superior text normalization system.
Subhojeet Pramanik, Aman Hussain
arXiv:1806.00044 · cs.CL · submitted May 31, 2018 · updated Apr 3, 2019
abstract · pdf · html · 9 pages, 10 tables, 3 figures
The write up on that (from Google, who organized it and provided the data) was really interesting: http://blog.kaggle.com/2018/02/07/a-brief-summary-of-the-kag...