about
Are we at “imagenet moment” for ASR? (arxiv.org)
2 points by option on May 12, 2020 | hide | past | pdf | discuss on HN

In plain words: An English speech-recognition system keeps learning from new accents, other languages, or specialized speech instead of starting its training over. It transcribed more accurately than systems trained from scratch, and bigger starting models stayed best even with only a little new speech.

Abstract · Cross-Language Transfer Learning, Continuous Learning, and Domain Adaptation for End-to-End Automatic Speech Recognition

In this paper, we demonstrate the efficacy of transfer learning and continuous learning for various automatic speech recognition (ASR) tasks. We start with a pre-trained English ASR model and show that transfer learning can be effectively and easily performed on: (1) different English accents, (2) different languages (German, Spanish and Russian) and (3) application-specific domains. Our experiments demonstrate that in all three cases, transfer learning from a good base model has higher accuracy than a model trained from scratch. It is preferred to fine-tune large models than small pre-trained models, even if the dataset for fine-tuning is small. Moreover, transfer learning significantly speeds up convergence for both very small and very large target datasets.

Jocelyn Huang, Oleksii Kuchaiev, Patrick O'Neill, Vitaly Lavrukhin, Jason Li, Adriana Flores, Georg Kucsko, Boris Ginsburg
arXiv:2005.04290 · eess.AS · submitted May 8, 2020
abstract · pdf · html

add comment on HN