about
Optimizing Speech Recognition for the Edge (arxiv.org)
2 points by sel1 on Oct 1, 2019 | hide | past | pdf | discuss on HN

In plain words: They start with a standard speech-to-text network of memory layers, then swap in cheaper layers and shrink it by cutting connections and storing numbers with less precision. It runs on a phone and is ten times smaller than the unoptimized model, with high accuracy.

Abstract · Optimizing Speech Recognition For The Edge

While most deployed speech recognition systems today still run on servers, we are in the midst of a transition towards deployments on edge devices. This leap to the edge is powered by the progression from traditional speech recognition pipelines to end-to-end (E2E) neural architectures, and the parallel development of more efficient neural network topologies and optimization techniques. Thus, we are now able to create highly accurate speech recognizers that are both small and fast enough to execute on typical mobile devices. In this paper, we begin with a baseline RNN-Transducer architecture comprised of Long Short-Term Memory (LSTM) layers. We then experiment with a variety of more computationally efficient layer types, as well as apply optimization techniques like neural connection pruning and parameter quantization to construct a small, high quality, on-device speech recognizer that is an order of magnitude smaller than the baseline system without any optimizations.

Yuan Shangguan, Jian Li, Qiao Liang, Raziel Alvarez, Ian McGraw
arXiv:1909.12408 · cs.CL, cs.LG, eess.AS · submitted Sep 26, 2019 · updated Feb 7, 2020
abstract · pdf · html

add comment on HN
Also discussed: Oct 2019 (4 points, 0 comments)