In plain words: An open-source toolkit for training speech recognizers that turn audio straight into text, built to spread training across many machines and to include a fast parallel decoder that blends in word predictions. It matched the best accuracy on standard speech benchmarks without extra training data and decoded 4–11 times faster than similar toolkits.
Abstract · Espresso: A Fast End-to-end Neural Speech Recognition Toolkit
We present Espresso, an open-source, modular, extensible end-to-end neural automatic speech recognition (ASR) toolkit based on the deep learning library PyTorch and the popular neural machine translation toolkit fairseq. Espresso supports distributed training across GPUs and computing nodes, and features various decoding approaches commonly employed in ASR, including look-ahead word-based language model fusion, for which a fast, parallelized decoder is implemented. Espresso achieves state-of-the-art ASR performance on the WSJ, LibriSpeech, and Switchboard data sets among other end-to-end systems without data augmentation, and is 4--11x faster for decoding than similar systems (e.g. ESPnet).
Yiming Wang, Tongfei Chen, Hainan Xu, Shuoyang Ding, Hang Lv, Yiwen Shao, Nanyun Peng, Lei Xie, Shinji Watanabe, Sanjeev Khudanpur
arXiv:1909.08723 · cs.CL, cs.SD, eess.AS · submitted Sep 18, 2019 · updated Oct 15, 2019
abstract · pdf · html · Accepted to ASRU 2019
Disadvantages are:
1) A bit disorgranized codebase with directly imported fairseq
2) No online decoding in design which is a must for real-world applications.
Other toolkits:
1) ESPnet - crazy dual chainer/pytorch backend, pretty slow from beginning, otherwise good.
2) Mozilla DeepSpeech - very lightweight technology, no real accuracy and speed.
3) nvidia/NEMO - potentially good performance from GPU experts, but not clear how it will develop in the future
4) speechbrain - just announced, no real code
5) facebook/wav2letter - C++ codebase, not within general NN community
6) tensoflow/lingvo - a playground for Google guys, who uses tensorflow these days?
7) kaldi - good old one (if 7 years is old for you), still has very important features others do not have (semi-supervised learning, long alignment). But no Pytorch again, not very attractive for general NN community.
8) didi/delta - did anyone try it at all?
9) PaddlePaddle/DeepSpeech - very old technology too, but Baidu releases very good models trained on their proprietary data