about
Surprisal-Triggered Conditional Computation with Neural Networks (arxiv.org)
1 point by JamilD on Jun 4, 2020 | hide | past | pdf | discuss on HN

In plain words: A predictor scores how surprising each new input is; surprising, hard ones go to a big network, easy ones to a small fast one. On two speech recognition tasks it matched always using the big network while doing 15% less computing work.

Abstract

Autoregressive neural network models have been used successfully for sequence generation, feature extraction, and hypothesis scoring. This paper presents yet another use for these models: allocating more computation to more difficult inputs. In our model, an autoregressive model is used both to extract features and to predict observations in a stream of input observations. The surprisal of the input, measured as the negative log-likelihood of the current observation according to the autoregressive model, is used as a measure of input difficulty. This in turn determines whether a small, fast network, or a big, slow network, is used. Experiments on two speech recognition tasks show that our model can match the performance of a baseline in which the big network is always used with 15% fewer FLOPs.

Loren Lugosch, Derek Nowrouzezahrai, Brett H. Meyer
arXiv:2006.01659 · cs.LG, stat.ML · submitted Jun 2, 2020
abstract · pdf · html

add comment on HN