In plain words: Re-testing 18 recent neural recommendation algorithms showed whether published results could be reproduced and how they compared with simple baselines. Only 7 were reproducible, and 6 were often beaten by simple nearest-neighbor or graph tricks; the last barely beat a well-tuned linear method.
Abstract · Are We Really Making Much Progress? A Worrying Analysis of Recent Neural Recommendation Approaches
Deep learning techniques have become the method of choice for researchers working on algorithmic aspects of recommender systems. With the strongly increased interest in machine learning in general, it has, as a result, become difficult to keep track of what represents the state-of-the-art at the moment, e.g., for top-n recommendation tasks. At the same time, several recent publications point out problems in today's research practice in applied machine learning, e.g., in terms of the reproducibility of the results or the choice of the baselines when proposing new models. In this work, we report the results of a systematic analysis of algorithmic proposals for top-n recommendation tasks. Specifically, we considered 18 algorithms that were presented at top-level research conferences in the last years. Only 7 of them could be reproduced with reasonable effort. For these methods, it however turned out that 6 of them can often be outperformed with comparably simple heuristic methods, e.g., based on nearest-neighbor or graph-based techniques. The remaining one clearly outperformed the baselines but did not consistently outperform a well-tuned non-neural linear ranking method. Overall, our work sheds light on a number of potential problems in today's machine learning scholarship and calls for improved scientific practices in this area. Source code of our experiments and full results are available at: https://github.com/MaurizioFD/RecSys2019_DeepLearning_Evaluation.
Maurizio Ferrari Dacrema, Paolo Cremonesi, Dietmar Jannach
arXiv:1907.06902 · cs.IR, cs.LG, cs.NE · submitted Jul 16, 2019 · updated Aug 16, 2019
abstract · pdf · Source code available at: https://github.com/MaurizioFD/RecSys2019_DeepLearning_Evaluation
Does this mean that the published results are garbage? No! But what does it mean?
First, don't miss the main claim: less than half the published papers have sufficient code and data available for replication. While one might argue exactly what the standards for availability should be, if journals and reviewers were to demand it, this number can definitely be improved. Greater availability of data and code makes for better science.
Second, my guess is that it means that paper authors are using excess "puffery" in an attempt to improve the chances that their papers are accepted for publication. They've decided that using a weak baseline and claiming a 50% improvement is more likely to result in publication than using a realistic baseline and claiming a 0-5% improvement. Unfortunately, they are probably right.
The useful change (in my opinion) would be to change the standards by which papers are judged to be publishable. Reviewers should push back against weak baselines, but be more accepting of results that make modest claims. You can still have a useful paper even if the degree of improvement is small --- or even nonexistent. Maybe it will inspire someone else, maybe it's useful in an ensemble, or maybe it applies to a case where the baseline wouldn't. It's the puffery that's the problem, not the publication. Creating a system that rewards authors who more honestly evaluate their work would be an overall win.