In plain words: The study re-checked the experiments behind four years of papers that teach computers to judge how similar two images are. Many used flawed tests, and once corrected, gains over decade-old methods were marginal rather than huge.
Abstract · A Metric Learning Reality Check
Deep metric learning papers from the past four years have consistently claimed great advances in accuracy, often more than doubling the performance of decade-old methods. In this paper, we take a closer look at the field to see if this is actually true. We find flaws in the experimental methodology of numerous metric learning papers, and show that the actual improvements over time have been marginal at best.
Kevin Musgrave, Serge Belongie, Ser-Nam Lim
arXiv:2003.08505 · cs.CV · submitted Mar 18, 2020 · updated Sep 16, 2020
abstract · pdf · html · Visit https://www.github.com/KevinMusgrave/powerful-benchmarker for supplementary material, including the source code, configuration files, log files, and interactive bayesian optimization plots