In plain words: A trick that finds how an equation's answer shifts when its settings change, once only for smooth equations, now works for noisy ones: solve a second noisy equation and store the random bits. It fits neural dynamics on 50-dimensional motion data competitively using constant memory.
Abstract
The adjoint sensitivity method scalably computes gradients of solutions to ordinary differential equations. We generalize this method to stochastic differential equations, allowing time-efficient and constant-memory computation of gradients with high-order adaptive solvers. Specifically, we derive a stochastic differential equation whose solution is the gradient, a memory-efficient algorithm for caching noise, and conditions under which numerical solutions converge. In addition, we combine our method with gradient-based stochastic variational inference for latent stochastic differential equations. We use our method to fit stochastic dynamics defined by neural networks, achieving competitive performance on a 50-dimensional motion capture dataset.
Xuechen Li, Ting-Kam Leonard Wong, Ricky T. Q. Chen, David Duvenaud
arXiv:2001.01328 · cs.LG, math.NA, stat.ML · submitted Jan 5, 2020 · updated Oct 18, 2020
abstract · pdf · html · AISTATS 2020; 25 pages, 6 figures in main text; clarify notation in appendix