In plain words: Re-testing 18 recent neural recommendation algorithms showed whether published results could be reproduced and how they compared with simple baselines. Only 7 were reproducible, and 6 were often beaten by simple nearest-neighbor or graph tricks; the last barely beat a well-tuned linear method.
Abstract · Are We Really Making Much Progress? A Worrying Analysis of Recent Neural Recommendation Approaches
Deep learning techniques have become the method of choice for researchers working on algorithmic aspects of recommender systems. With the strongly increased interest in machine learning in general, it has, as a result, become difficult to keep track of what represents the state-of-the-art at the moment, e.g., for top-n recommendation tasks. At the same time, several recent publications point out problems in today's research practice in applied machine learning, e.g., in terms of the reproducibility of the results or the choice of the baselines when proposing new models. In this work, we report the results of a systematic analysis of algorithmic proposals for top-n recommendation tasks. Specifically, we considered 18 algorithms that were presented at top-level research conferences in the last years. Only 7 of them could be reproduced with reasonable effort. For these methods, it however turned out that 6 of them can often be outperformed with comparably simple heuristic methods, e.g., based on nearest-neighbor or graph-based techniques. The remaining one clearly outperformed the baselines but did not consistently outperform a well-tuned non-neural linear ranking method. Overall, our work sheds light on a number of potential problems in today's machine learning scholarship and calls for improved scientific practices in this area. Source code of our experiments and full results are available at: https://github.com/MaurizioFD/RecSys2019_DeepLearning_Evaluation.
Maurizio Ferrari Dacrema, Paolo Cremonesi, Dietmar Jannach
arXiv:1907.06902 · cs.IR, cs.LG, cs.NE · submitted Jul 16, 2019 · updated Aug 16, 2019
abstract · pdf · Source code available at: https://github.com/MaurizioFD/RecSys2019_DeepLearning_Evaluation
1) Half the papers couldn't be reproduced on a technical level. Publish your code and your data, people!
2) Most of these papers uses "weak baselines" so they can show some kind of improvement and get their paper published. I'm conflicted about this because if we require every paper to beat state-of-the-art, we'd (collectively, as the entire discipline) be lucky to publish one paper a year. From one point of view, these papers actually represent a form of publishing negative results - we tried this and it didn't work - which isn't a bad thing. But the biased way its presented makes it harder to separate the wheat from the chaff.
3) It's not obvious that we're going to squeeze any more business value from this particular stone. Sometimes all the useful information in a dataset can be found with a fairly simple algorithm. Not everything benefits from a more complex representation, and sometimes you can't fix that with regularization or more data. Sometimes you just have to use the simple model and accept that it captures all the signal that's available and everything else is noise.