In plain words: A team tried to rebuild 255 papers from 1984 to 2017 using only the paper's text, never the authors' code, and logged what helped or blocked each attempt. The analysis looks for which paper features predict success, because sharing code alone does not guarantee reproducibility.
Abstract
What makes a paper independently reproducible? Debates on reproducibility center around intuition or assumptions but lack empirical results. Our field focuses on releasing code, which is important, but is not sufficient for determining reproducibility. We take the first step toward a quantifiable answer by manually attempting to implement 255 papers published from 1984 until 2017, recording features of each paper, and performing statistical analysis of the results. For each paper, we did not look at the authors code, if released, in order to prevent bias toward discrepancies between code and paper.
Edward Raff
arXiv:1909.06674 · cs.LG, cs.AI, cs.DL, stat.ML · submitted Sep 14, 2019
abstract · pdf · html · to appear in Proc. Neural Information Processing Systems (NeurIPS), 2019