In plain words: Many deep learning failures share one cause: the model grabs an easy rule that scores well on standard tests but breaks in the real world. It lays out guidance for checking what a model really learned and testing it in harder, more realistic conditions.
Abstract
Deep learning has triggered the current rise of artificial intelligence and is the workhorse of today's machine intelligence. Numerous success stories have rapidly spread all over science, industry and society, but its limitations have only recently come into focus. In this perspective we seek to distill how many of deep learning's problems can be seen as different symptoms of the same underlying problem: shortcut learning. Shortcuts are decision rules that perform well on standard benchmarks but fail to transfer to more challenging testing conditions, such as real-world scenarios. Related issues are known in Comparative Psychology, Education and Linguistics, suggesting that shortcut learning may be a common characteristic of learning systems, biological and artificial alike. Based on these observations, we develop a set of recommendations for model interpretation and benchmarking, highlighting recent advances in machine learning to improve robustness and transferability from the lab to real-world applications.
Robert Geirhos, Jörn-Henrik Jacobsen, Claudio Michaelis, Richard Zemel, Wieland Brendel, Matthias Bethge, Felix A. Wichmann
arXiv:2004.07780 · cs.CV, cs.AI, cs.LG, q-bio.NC · submitted Apr 16, 2020 · updated Nov 21, 2023
abstract · pdf · html · perspective article published at Nature Machine Intelligence (https://doi.org/10.1038/s42256-020-00257-z)
Odd OT question: how did you end up including annotations in the bibliography? I found that super useful but it's incredibly uncommon in what I normally read.