In plain words: Volunteers record sentences and check each other's clips, making a free, huge transcribed audio collection for training speech recognition in many languages. Adapting an English speech-to-text model to twelve languages cut character errors by about 6 points on average, the first published results for most.
Abstract
The Common Voice corpus is a massively-multilingual collection of transcribed speech intended for speech technology research and development. Common Voice is designed for Automatic Speech Recognition purposes but can be useful in other domains (e.g. language identification). To achieve scale and sustainability, the Common Voice project employs crowdsourcing for both data collection and data validation. The most recent release includes 29 languages, and as of November 2019 there are a total of 38 languages collecting data. Over 50,000 individuals have participated so far, resulting in 2,500 hours of collected audio. To our knowledge this is the largest audio corpus in the public domain for speech recognition, both in terms of number of hours and number of languages. As an example use case for Common Voice, we present speech recognition experiments using Mozilla's DeepSpeech Speech-to-Text toolkit. By applying transfer learning from a source English model, we find an average Character Error Rate improvement of 5.99 +/- 5.48 for twelve target languages (German, French, Italian, Turkish, Catalan, Slovenian, Welsh, Irish, Breton, Tatar, Chuvash, and Kabyle). For most of these languages, these are the first ever published results on end-to-end Automatic Speech Recognition.
Rosana Ardila, Megan Branson, Kelly Davis, Michael Henretty, Michael Kohler, Josh Meyer, Reuben Morais, Lindsay Saunders, Francis M. Tyers, Gregor Weber
arXiv:1912.06670 · cs.CL, cs.LG · submitted Dec 13, 2019 · updated Mar 5, 2020
abstract · pdf · html · Accepted to LREC 2020
There is certainly an aspect, as M-AILABS sources their data mainly from the librivox project, which is community-driven.
To give ballpark numbers on what it would have cost if you had had to pay people for providing the data instead of getting it for free:
It's low skilled labour so you'll likely find people to do it slightly above minimum wage. Let's take Germany as I'm most familiar with its rules and because that's where Common Voice is headquartered. Minimum wage here will be €9.35 / hr starting on Jan 1st. Let's say you pay them €11. There are various Arbeitgeberanteile which you have to pay as well. Let's say your per-employee expense would amount to €15/hour. Let's assume you can verify and record at 70% efficiency and you use two people to verify. Then you need to expend 4.29 employee-hours per final result hour.
You couldn't just get German at this rate: Berlin is one of the towns with the largest language diversities in Germany.
This would give you a price of €64.35 per result hour. You'd have to pay €64k for 1000 hours of validated training data, and 128k for the 2 thousand hour figure of currently achieved data.
These €128k are probably on the same order of magnitude that Mozilla pays for the project (employee time to design, build, and run it), and if the project scales it will look even better. From a business POV, going open source was thus a great idea.
To put the 2000 hours into comparison, the deepspeech 2 paper [1] used 10k hour datasets (per language, while the 2k hours are distributed amongst multiple languages) [1]. Record holder is probably Amazon with 1 million hours (although it's unlabeled) [2].
It's possible though that future breakthroughs will remove the need for tons of training data. So even if the restricted amount of training data can't create practical models in niche languages for now, it might very well be able to in the future.
[1]: https://arxiv.org/pdf/1512.02595.pdf
[1]: https://arxiv.org/pdf/1904.01624.pdf