In plain words: A course extends calculus to functions of matrices and vectors, showing how to differentiate matrix inverses and other complicated calculations. It teaches reverse-mode differentiation, which sweeps backward through a computation to get every derivative at once instead of one at a time.
Abstract · Matrix Calculus (for Machine Learning and Beyond)
This course, intended for undergraduates familiar with elementary calculus and linear algebra, introduces the extension of differential calculus to functions on more general vector spaces, such as functions that take as input a matrix and return a matrix inverse or factorization, derivatives of ODE solutions, and even stochastic derivatives of random functions. It emphasizes practical computational applications, such as large-scale optimization and machine learning, where derivatives must be re-imagined in order to be propagated through complicated calculations. The class also discusses efficiency concerns leading to "adjoint" or "reverse-mode" differentiation (a.k.a. "backpropagation"), and gives a gentle introduction to modern automatic differentiation (AD) techniques.
Paige Bright, Alan Edelman, Steven G. Johnson
arXiv:2501.14787 · math.HO, cs.LG, math.NA, stat.ML · submitted Jan 7, 2025
abstract · pdf · html · Lecture notes for the MIT short course 18.063 "Matrix Calculus", based on the class as taught in January 2023 (also available on MIT OpenCourseWare)
In a graduate numerical optimization class I took over a decade ago, the professor spent 10 minutes on the first day deriving some matrix calculus identity by working out the expressions for partial derivatives using simple calculus rules and a lot of manual labor. Then, as the class was winding up, he joked and said "just kidding, don't do that... here's how we can do this with a Taylor expansion", and proceeded to derive the same identity in what felt like 30 seconds.
Also, don't forget the Jacobian and gradient aren't the same thing!