about
The Life of a Dataset in Machine Learning Research (arxiv.org)
2 points by headalgorithm on Jan 26, 2023 | hide | past | pdf | 1 comment on HN

In plain words: They tracked which benchmark datasets machine learning papers used from 2015 to 2020 across different research areas. Each area came to rely on fewer datasets over time, even as papers increasingly borrowed datasets built for other tasks.

Abstract · Reduced, Reused and Recycled: The Life of a Dataset in Machine Learning Research

Benchmark datasets play a central role in the organization of machine learning research. They coordinate researchers around shared research problems and serve as a measure of progress towards shared goals. Despite the foundational role of benchmarking practices in this field, relatively little attention has been paid to the dynamics of benchmark dataset use and reuse, within or across machine learning subcommunities. In this paper, we dig into these dynamics. We study how dataset usage patterns differ across machine learning subcommunities and across time from 2015-2020. We find increasing concentration on fewer and fewer datasets within task communities, significant adoption of datasets from other tasks, and concentration across the field on datasets that have been introduced by researchers situated within a small number of elite institutions. Our results have implications for scientific evaluation, AI ethics, and equity/access within the field.

Bernard Koch, Emily Denton, Alex Hanna, Jacob G. Foster
arXiv:2112.01716 · cs.LG, cs.CL, cs.CV, cs.CY, stat.ML · submitted Dec 3, 2021
abstract · pdf · html · 35th Conference on Neural Information Processing Systems (NeurIPS 2021), Sydney, Australia

add comment on HN
Also discussed: Dec 2021 (1 point, 0 comments)

Abstract:

Benchmark datasets play a central role in the organization of machine learning research. They coordinate researchers around shared research problems and serve as a measure of progress towards shared goals. Despite the foundational role of benchmarking practices in this field, relatively little attention has been paid to the dynamics of benchmark dataset use and reuse, within or across machine learning subcommunities.

In this paper, we dig into these dynamics. We study how dataset usage patterns differ across machine learning subcommunities and across time from 2015-2020. We find increasing concentration on fewer and fewer datasets within task communities, significant adoption of datasets from other tasks, and concentration across the field on datasets that have been introduced by researchers situated within a small number of elite institutions.

Our results have implications for scientific evaluation, AI ethics, and equity/access within the field.