In plain words: Sorting and ranking have sharp corners or flat spots, so learning systems can't compute useful gradients through them. The new operators run as fast as ordinary sorting and give exact answers and slopes, making them about ten times faster than earlier smooth versions.
Abstract
The sorting operation is one of the most commonly used building blocks in computer programming. In machine learning, it is often used for robust statistics. However, seen as a function, it is piecewise linear and as a result includes many kinks where it is non-differentiable. More problematic is the related ranking operator, often used for order statistics and ranking metrics. It is a piecewise constant function, meaning that its derivatives are null or undefined. While numerous works have proposed differentiable proxies to sorting and ranking, they do not achieve the $O(n \log n)$ time complexity one would expect from sorting and ranking operations. In this paper, we propose the first differentiable sorting and ranking operators with $O(n \log n)$ time and $O(n)$ space complexity. Our proposal in addition enjoys exact computation and differentiation. We achieve this feat by constructing differentiable operators as projections onto the permutahedron, the convex hull of permutations, and using a reduction to isotonic optimization. Empirically, we confirm that our approach is an order of magnitude faster than existing approaches and showcase two novel applications: differentiable Spearman's rank correlation coefficient and least trimmed squares.
Mathieu Blondel, Olivier Teboul, Quentin Berthet, Josip Djolonga
arXiv:2002.08871 · stat.ML, cs.LG · submitted Feb 20, 2020 · updated Jun 29, 2020
abstract · pdf · html · In proceedings of ICML 2020
https://en.wikipedia.org/wiki/Doubly_stochastic_matrix
which is better explained here:
https://cs.stackexchange.com/questions/4805/sorting-as-a-lin...
where the sorting operation is defined as a linear program. Evidently, this has been known for at least half a century. That said, if a solution to a linear program can be found in a way that's differentiable, this means that the operation of sorting can be found to be differentiable as well. This appears to be trick in the paper and they appear to have a relatively fast way to compute this solution as well, which I think is interesting.