In plain words: A small network spots spoken keywords by combining layers that catch local sound patterns with layers that track longer context, instead of big speech-recognition networks. With about 230,000 learned values, it hit 97.71% accuracy at 0.5 false alarms per hour in noisy audio.
Abstract · Convolutional Recurrent Neural Networks for Small-Footprint Keyword Spotting
Keyword spotting (KWS) constitutes a major component of human-technology interfaces. Maximizing the detection accuracy at a low false alarm (FA) rate, while minimizing the footprint size, latency and complexity are the goals for KWS. Towards achieving them, we study Convolutional Recurrent Neural Networks (CRNNs). Inspired by large-scale state-of-the-art speech recognition systems, we combine the strengths of convolutional layers and recurrent layers to exploit local structure and long-range context. We analyze the effect of architecture parameters, and propose training strategies to improve performance. With only ~230k parameters, our CRNN model yields acceptably low latency, and achieves 97.71% accuracy at 0.5 FA/hour for 5 dB signal-to-noise ratio.
Sercan O. Arik, Markus Kliegl, Rewon Child, Joel Hestness, Andrew Gibiansky, Chris Fougner, Ryan Prenger, Adam Coates
arXiv:1703.05390 · cs.CL, cs.AI, cs.LG · submitted Mar 15, 2017 · updated Jul 4, 2017
abstract · pdf · Accepted to Interspeech 2017