about
Dynamic Automatic Differentiation of GPU Broadcast Kernels (arxiv.org)
5 points by jrevels on Oct 23, 2018 | hide | past | pdf | 1 comment on HN

In plain words: When one operation is applied to many values at once, derivatives come from pushing changes forward through that step, not backward. It skips backward replay and lets the step be fused even when code branches on data; a Julia test beat backward-mode versions.

Abstract

We show how forward-mode automatic differentiation (AD) can be employed within larger reverse-mode computations to dynamically differentiate broadcast operations in a GPU-friendly manner. Our technique fully exploits the broadcast Jacobian's inherent sparsity structure, and unlike a pure reverse-mode approach, this "mixed-mode" approach does not require a backwards pass over the broadcasted operation's subgraph, obviating the need for several reverse-mode-specific programmability restrictions on user-authored broadcast operations. Most notably, this approach allows broadcast fusion in primal code despite the presence of data-dependent control flow. We discuss an experiment in which a Julia implementation of our technique outperformed pure reverse-mode TensorFlow and Julia implementations for differentiating through broadcast operations within an HM-LSTM cell update calculation.

Jarrett Revels, Tim Besard, Valentin Churavy, Bjorn De Sutter, Juan Pablo Vielma
arXiv:1810.08297 · cs.MS · submitted Oct 18, 2018 · updated Oct 24, 2018
abstract · pdf · html

add comment on HN

One of the author here. Working on this paper was fun, one of the most interesting lessons I learned was that divergence on GPUs has dramatically fallen in cost on modern architectures.