about
Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages (arxiv.org)
2 points by pallas_athena on Mar 7, 2023 | hide | past | pdf | discuss on HN

In plain words: One model learns from 12 million hours of unlabeled speech in 300+ languages, then trains on a small labeled set to transcribe 100+ languages. It matched or beat a widely used model trained on seven times more labeled audio, on familiar and unfamiliar speech.

Abstract

We introduce the Universal Speech Model (USM), a single large model that performs automatic speech recognition (ASR) across 100+ languages. This is achieved by pre-training the encoder of the model on a large unlabeled multilingual dataset of 12 million (M) hours spanning over 300 languages, and fine-tuning on a smaller labeled dataset. We use multilingual pre-training with random-projection quantization and speech-text modality matching to achieve state-of-the-art performance on downstream multilingual ASR and speech-to-text translation tasks. We also demonstrate that despite using a labeled training set 1/7-th the size of that used for the Whisper model, our model exhibits comparable or better performance on both in-domain and out-of-domain speech recognition tasks across many languages.

Yu Zhang, Wei Han, James Qin, Yongqiang Wang, Ankur Bapna, Zhehuai Chen, Nanxin Chen, Bo Li, Vera Axelrod, Gary Wang, Zhong Meng, Ke Hu, et al.
arXiv:2303.01037 · cs.CL, cs.SD, eess.AS · submitted Mar 2, 2023 · updated Sep 25, 2023
abstract · pdf · html · 20 pages, 7 figures, 8 tables

add comment on HN