In plain words: Swapping which tasks a benchmark includes can flip which machine-learning method looks best, even when the methods themselves have not changed. Rankings shift significantly across common benchmark setups, so a win may reflect task choice rather than real superiority.
Abstract
The world of empirical machine learning (ML) strongly relies on benchmarks in order to determine the relative effectiveness of different algorithms and methods. This paper proposes the notion of "a benchmark lottery" that describes the overall fragility of the ML benchmarking process. The benchmark lottery postulates that many factors, other than fundamental algorithmic superiority, may lead to a method being perceived as superior. On multiple benchmark setups that are prevalent in the ML community, we show that the relative performance of algorithms may be altered significantly simply by choosing different benchmark tasks, highlighting the fragility of the current paradigms and potential fallacious interpretation derived from benchmarking ML methods. Given that every benchmark makes a statement about what it perceives to be important, we argue that this might lead to biased progress in the community. We discuss the implications of the observed phenomena and provide recommendations on mitigating them using multiple machine learning domains and communities as use cases, including natural language processing, computer vision, information retrieval, recommender systems, and reinforcement learning.
Mostafa Dehghani, Yi Tay, Alexey A. Gritsenko, Zhe Zhao, Neil Houlsby, Fernando Diaz, Donald Metzler, Oriol Vinyals
arXiv:2107.07002 · cs.LG, cs.AI, cs.CL, cs.CV, cs.IR · submitted Jul 14, 2021
abstract · pdf · html
But what is the "original problem" and how do we measure progress towards solving it? Obviously there's not just one such problem - each community has a few of its own.
But in general, the reason that we waste so much time and effort on benchmarks in AI research (and that's AI in general, not just machine learning this time) is because nobody can really answer this fundamental question: how do we measure the progress of AI research?
And that in turn is because AI research is not guided by a scientific theory: an epistemic object that can explain current and past observations according to current and past knowledge, and make predictions of future observations. We do not have such a theory of artificial intelligence. Therefore, we do not know what we are doing, we do not know where we are going and we do not even know where we are.
This is the sad, sad state of AI research. If AI research has been reduced, time and again, to a spectacle, a race to the bottom of pointless benchmarks, that's because AI research has never stopped to take its bearings, figure out its goals (there are no commonly accepted goals of AI research) and establish itself as a science, with a theory - rather than a constantly shifting trip from demonstration to demonstration. 70 years of demonstrations!
I think the paper above manages to go on about benchmarks for 34 pages and still miss the real limitation of empirical-only evaluations in a field without a theoretical basis. That no matter what benchmarks you choose and how, without a theoretical basis, you'll never know what you're doing.