In plain words: A network learns sound patterns from unlabeled audio by spotting which parts are real and which were altered, and uses them to improve speech recognition. With just a few hours of transcribed speech, it cut word errors up to 36% versus a strong baseline.
Abstract · wav2vec: Unsupervised Pre-training for Speech Recognition
We explore unsupervised pre-training for speech recognition by learning representations of raw audio. wav2vec is trained on large amounts of unlabeled audio data and the resulting representations are then used to improve acoustic model training. We pre-train a simple multi-layer convolutional neural network optimized via a noise contrastive binary classification task. Our experiments on WSJ reduce WER of a strong character-based log-mel filterbank baseline by up to 36% when only a few hours of transcribed data is available. Our approach achieves 2.43% WER on the nov92 test set. This outperforms Deep Speech 2, the best reported character-based system in the literature while using two orders of magnitude less labeled training data.
Steffen Schneider, Alexei Baevski, Ronan Collobert, Michael Auli
arXiv:1904.05862 · cs.CL · submitted Apr 11, 2019 · updated Sep 11, 2019
abstract · pdf · html
The approach seems to provide significant accuracy boost when the supervised training set available is small, e.g., less than 10 hours. The relative improvement is modest over baseline supervised model trained on 10s of hours of transcribed audio. The trends indicate that the improvement is probably minimal when 100s of hours of supervised training data is available.
The authors report improvements over Deepspeech on certain datasets. Deepspeech uses a 5-gram language model. The proposed model has significantly lower performance (albeit on a smaller supervised training set) when it also uses an n-gram-based language model. Improvements over Deepspeech are shown when convolutional language models are used. Hence, it is possible that the improvements over Deepspeech are contributed mainly by the use of convolutional language models. Comparing with Deepspeech+conv language model will provide a better apple-to-apple comparison of the proposed unsupervised pre-trained acoustic model.
The gains also seem to have diminishing returns as the number of hours of unsupervised training data increases (improvement is marginal even with 10x increase of unsupervised training data).