In plain words: A book reviews the core ideas of differentiable programming, where programs—including loops and branching—can be differentiated so their settings are tuned gradually to cut errors. It shows that making programs differentiable also spreads a probability over how they run, so you can measure uncertainty.
Abstract
Artificial intelligence has recently experienced remarkable advances, fueled by large models, vast datasets, accelerated hardware, and, last but not least, the transformative power of differentiable programming. This new programming paradigm enables end-to-end differentiation of complex computer programs (including those with control flows and data structures), making gradient-based optimization of program parameters possible. As an emerging paradigm, differentiable programming builds upon several areas of computer science and applied mathematics, including automatic differentiation, graphical models, optimization and statistics. This book presents a comprehensive review of the fundamental concepts useful for differentiable programming. We adopt two main perspectives, that of optimization and that of probability, with clear analogies between the two. Differentiable programming is not merely the differentiation of programs, but also the thoughtful design of programs intended for differentiation. By making programs differentiable, we inherently introduce probability distributions over their execution, providing a means to quantify the uncertainty associated with program outputs.
Mathieu Blondel, Vincent Roulet
arXiv:2403.14606 · cs.LG, cs.AI, cs.PL · submitted Mar 21, 2024 · updated Aug 3, 2026
abstract · pdf · html · Draft version 4
Every element in the dual numbers is of the form a + bh, and in fact the entire ring can be turned into a totally ordered ring in a very natural way: simply declare h < r for any real r > 0. In essence, we are saying h is an infinitesimal - so small that its square is 0. So we have a non-Archimedean ring with infinitesimals - the smallest such ring extending the real numbers.
Why is this so important? Well, if you have some function f which can be extended to the dual number plane - which many can, similar to the complex plane - we have
f(x+h) = f(x) + f'(x)h
Which is little more than restating the usual definition of the derivative: f'(x) = (f(x+h) - f(x))/h
For instance, suppose we have f(x) = 2x² - 3x + 1, then
f(x+h) = 2(x+h)² - 3(x+h) + 1 = 2(x² + 2xh + h²) - 3(x+h) + 1 = (2x² - 3x + 1) + (4x - 3)h
Where the last step just involves rearranging terms and canceling out the h² = 0 term. Note that the expression for the derivative we get, (4x-3), is correct, and magically computed itself straight from the properties of the algebra.
In short, just like creating i² = -1 revolutionized algebra, setting h² = 0 revolutionizes calculus. Most autodiff packages (such as Pytorch) use something not much more advanced than this, although there are optimizations to speed it up (e.g. reverse mode diff).