about
Deep Speech: Scaling up end-to-end speech recognition (arxiv.org)
63 points by lelf on Feb 14, 2015 | hide | past | pdf | 8 comments on HN

In plain words: A neural network turns speech into text, skipping the hand-built steps and phoneme dictionary traditional systems need, and learns from varied, artificially mixed audio to resist noise. It hit 16.0% error on a standard speech test, beating published results and commercial systems in noise.

Abstract

We present a state-of-the-art speech recognition system developed using end-to-end deep learning. Our architecture is significantly simpler than traditional speech systems, which rely on laboriously engineered processing pipelines; these traditional systems also tend to perform poorly when used in noisy environments. In contrast, our system does not need hand-designed components to model background noise, reverberation, or speaker variation, but instead directly learns a function that is robust to such effects. We do not need a phoneme dictionary, nor even the concept of a "phoneme." Key to our approach is a well-optimized RNN training system that uses multiple GPUs, as well as a set of novel data synthesis techniques that allow us to efficiently obtain a large amount of varied data for training. Our system, called Deep Speech, outperforms previously published results on the widely studied Switchboard Hub5'00, achieving 16.0% error on the full test set. Deep Speech also handles challenging noisy environments better than widely used, state-of-the-art commercial speech systems.

Awni Hannun, Carl Case, Jared Casper, Bryan Catanzaro, Greg Diamos, Erich Elsen, Ryan Prenger, Sanjeev Satheesh, Shubho Sengupta, Adam Coates, Andrew Y. Ng
arXiv:1412.5567 · cs.CL, cs.LG, cs.NE · submitted Dec 17, 2014 · updated Dec 19, 2014
abstract · pdf · html

add comment on HN
Also discussed: Aug 2018 (3 points, 0 comments) · Dec 2014 (4 points, 0 comments) · Dec 2014 (79 points, 22 comments)

I'll kick things off by linking to the comments from the last time this was posted: https://news.ycombinator.com/item?id=8769067.
>16.0% error on the full test set

Does anyone know, what was the error rate with previous approaches these days?

Look at Table 3 in the paper. Also, 16.5% error ;)
Oh thanks, and, wow!
Github repo that I can compile and try?
Providing the source would be a step to improve transparency and reproducibility (the text does not provide sufficient detail for even someone working in the field to reproduce what they did so that he would arrive at the same results); however, the more crucial thing is the data. Switchboard, Fisher, and WSJ are available (provided you have a few grand to spend), but they say they collected 5000 h of read speech from 9600 speakers.. That's a huge effort!
An alternative source of data that you can contribute to:

http://www.voxforge.org/

woo!