about
Recurrent Neural Nets for Speech Synthesis (arxiv.org)
88 points by Katydid on Jan 18, 2016 | hide | past | pdf | 11 comments on HN

In plain words: They studied a speech-making network whose gates control what it remembers, using visual analysis and tests to find which gate matters most. The trimmed-down version keeps the same sound quality with far fewer numbers to store, cutting the work to generate speech considerably.

Abstract · Investigating gated recurrent neural networks for speech synthesis

Recently, recurrent neural networks (RNNs) as powerful sequence models have re-emerged as a potential acoustic model for statistical parametric speech synthesis (SPSS). The long short-term memory (LSTM) architecture is particularly attractive because it addresses the vanishing gradient problem in standard RNNs, making them easier to train. Although recent studies have demonstrated that LSTMs can achieve significantly better performance on SPSS than deep feed-forward neural networks, little is known about why. Here we attempt to answer two questions: a) why do LSTMs work well as a sequence model for SPSS; b) which component (e.g., input gate, output gate, forget gate) is most important. We present a visual analysis alongside a series of experiments, resulting in a proposal for a simplified architecture. The simplified architecture has significantly fewer parameters than an LSTM, thus reducing generation complexity considerably without degrading quality.

Zhizheng Wu, Simon King
arXiv:1601.02539 · cs.CL, cs.NE · submitted Jan 11, 2016
abstract · pdf · html · Accepted by ICASSP 2016

add comment on HN

Went looking for audio samples, here's some from one of the researchers:

http://www.zhizheng.org/demo/is15_mte/demo.html http://www.zhizheng.org/demo/dnn_tts/demo.html

I thought this would be about text-to-speech applications, while this seems more like an encoder-decoder problem (make the network learn a pattern and then let it reproduce it). I'm wondering how long it is until we see working TTS based on LSTM RNNs.
Yeah, can someone explain the exact problem of "statistical parametric speech synthesis," since I can't find a general overview of the problem itself.
I'm a newbie to all this, but I can imagine it could be useful for speech compression.
This paper focuses on statistical parametric speech synthesis (SPSS). SPSS is only 1/2 of the text to speech problem.

SPSS is the problem of going from linguistic features, phenomes, etc, to speech audio. These features are more or less golden, either derived from the audio itself or hand-labeled. So things like tonality, cadence, emphasis on words is already encoded as features which is why these samples sound so good.

Deriving these features from pure text is very hard, and this failing is the main reason most text to speech systems sound so dull and tone-dead.

That being said, these results are seriously impressive, sounding very natural. Would love to see someone try and train an end-to-end system from pure text to speech. I think we'd see some big improvements like what Baidu has done for end-to-end speech to text.

The most interesting part of this paper is their simpler RNN structure than LSTM.
Slightly unrelated question, has there been any effort into hardware acceleration of such networks? How amenable are modern machine learning algorithms to hardware acceleration?
The GPU is pretty well optimized for the sort of operations an RNN needs.

There were a few efforts to make actual silicon neurons, plus the whole nueromorphic movement, but they were generally less than what people were expecting, slow, and difficult to interface with.

I've seen some work that attempts to recreate the "spiky" neural networks (e.g. neurons that fire when their inputs pass a threshold), intended to mimic the biochemistry of real neurons.

That work seems to spin their contribution as reducing the power required to evaluate the neural network though. If I recall correctly, the accuracy of those models for everyday tasks is typically much much lower than usual ANNs, and they're a pain to train. So, still not very common.

That is exactly what I made circa 2008. I used the izhekevich model for spiking. It was certainly faster on the GPU (2000x), but, yeah, getting the network to converge on anything was terrible. Debugging it was fun/awful though.

1:"Hey, do you see the first squiggle with the two fuzzes after."

2:"Next to Beaker's eyebrows?"

The low power work seems to have been aiming to be a rough filter, rather than a full system. Still fun to use.