about
Self-Teaching (Deep Neural) Networks (arxiv.org)
2 points by Anon84 on Sep 11, 2019 | hide | past | pdf | discuss on HN

In plain words: Lower layers are trained to copy the final layer's soft guesses through an extra loss, helping learning signals reach deeper layers and preventing overfitting. On speech recognition with 30,000 hours of audio, it beat the usual soft-label and confidence-penalty tricks in all four settings.

Abstract · Self-Teaching Networks

We propose self-teaching networks to improve the generalization capacity of deep neural networks. The idea is to generate soft supervision labels using the output layer for training the lower layers of the network. During the network training, we seek an auxiliary loss that drives the lower layer to mimic the behavior of the output layer. The connection between the two network layers through the auxiliary loss can help the gradient flow, which works similar to the residual networks. Furthermore, the auxiliary loss also works as a regularizer, which improves the generalization capacity of the network. We evaluated the self-teaching network with deep recurrent neural networks on speech recognition tasks, where we trained the acoustic model using 30 thousand hours of data. We tested the acoustic model using data collected from 4 scenarios. We show that the self-teaching network can achieve consistent improvements and outperform existing methods such as label smoothing and confidence penalization.

Liang Lu, Eric Sun, Yifan Gong
arXiv:1909.04157 · eess.AS, cs.CL, cs.LG, cs.SD, stat.ML · submitted Sep 9, 2019
abstract · pdf · html · 5 pages, Interspeech 2019

add comment on HN