In plain words: This review looks at what makes deep learning practical to compute: layered networks trained by repeatedly nudging weights using small random slices of data. It finds that fast linear algebra libraries, not the network design, are the key to handling huge datasets.
Abstract
In this article we review computational aspects of Deep Learning (DL). Deep learning uses network architectures consisting of hierarchical layers of latent variables to construct predictors for high-dimensional input-output models. Training a deep learning architecture is computationally intensive, and efficient linear algebra libraries is the key for training and inference. Stochastic gradient descent (SGD) optimization and batch sampling are used to learn from massive data sets.
Nicholas Polson, Vadim Sokolov
arXiv:1808.08618 · cs.LG, stat.CO, stat.ML · submitted Aug 26, 2018 · updated Aug 28, 2019
abstract · pdf · html