about
Neural Ordinary Differential Equations (arxiv.org)
3 points by kumaranvpl on Dec 13, 2018 | hide | past | pdf | discuss on HN

In plain words: Instead of stacking layers, this model sets a rule for how its hidden state changes and lets a math solver follow it to the answer. It uses the same memory no matter how deep, and can trade accuracy for speed.

Abstract

We introduce a new family of deep neural network models. Instead of specifying a discrete sequence of hidden layers, we parameterize the derivative of the hidden state using a neural network. The output of the network is computed using a black-box differential equation solver. These continuous-depth models have constant memory cost, adapt their evaluation strategy to each input, and can explicitly trade numerical precision for speed. We demonstrate these properties in continuous-depth residual networks and continuous-time latent variable models. We also construct continuous normalizing flows, a generative model that can train by maximum likelihood, without partitioning or ordering the data dimensions. For training, we show how to scalably backpropagate through any ODE solver, without access to its internal operations. This allows end-to-end training of ODEs within larger models.

Ricky T. Q. Chen, Yulia Rubanova, Jesse Bettencourt, David Duvenaud
arXiv:1806.07366 · cs.LG, cs.AI, stat.ML · submitted Jun 19, 2018 · updated Dec 14, 2019
abstract · pdf · html

add comment on HN
Also discussed: Dec 2018 (240 points, 60 comments)