In plain words: A system learns to recognize new spoken commands from just a few examples, while keeping a set of already-known commands in mind instead of treating every command as brand new. It beat standard training on many labeled examples and the usual few-shot learning approach.
Abstract · An Investigation of Few-Shot Learning in Spoken Term Classification
In this paper, we investigate the feasibility of applying few-shot learning algorithms to a speech task. We formulate a user-defined scenario of spoken term classification as a few-shot learning problem. In most few-shot learning studies, it is assumed that all the N classes are new in a N-way problem. We suggest that this assumption can be relaxed and define a N+M-way problem where N and M are the number of new classes and fixed classes respectively. We propose a modification to the Model-Agnostic Meta-Learning (MAML) algorithm to solve the problem. Experiments on the Google Speech Commands dataset show that our approach outperforms the conventional supervised learning approach and the original MAML.
Yangbin Chen, Tom Ko, Lifeng Shang, Xiao Chen, Xin Jiang, Qing Li
arXiv:1812.10233 · cs.CL, cs.IR · submitted Dec 26, 2018 · updated Sep 14, 2020
abstract · pdf · html · Accepted by INTERSPEECH 2020