about
Our datasets are flawed. ImageNet has an error rate of ~5.8% (arxiv.org)
23 points by chirau on Oct 28, 2021 | hide | past | pdf | 4 comments on HN

In plain words: A study checked the test sets of 10 vision, text, and audio datasets for wrong labels, using a scoring tool followed by human checks. At least 3.3% of labels were wrong on average, and fixing them can flip rankings so smaller models beat bigger ones.

Abstract · Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks

We identify label errors in the test sets of 10 of the most commonly-used computer vision, natural language, and audio datasets, and subsequently study the potential for these label errors to affect benchmark results. Errors in test sets are numerous and widespread: we estimate an average of at least 3.3% errors across the 10 datasets, where for example label errors comprise at least 6% of the ImageNet validation set. Putative label errors are identified using confident learning algorithms and then human-validated via crowdsourcing (51% of the algorithmically-flagged candidates are indeed erroneously labeled, on average across the datasets). Traditionally, machine learning practitioners choose which model to deploy based on test accuracy - our findings advise caution here, proposing that judging models over correctly labeled test sets may be more useful, especially for noisy real-world datasets. Surprisingly, we find that lower capacity models may be practically more useful than higher capacity models in real-world datasets with high proportions of erroneously labeled data. For example, on ImageNet with corrected labels: ResNet-18 outperforms ResNet-50 if the prevalence of originally mislabeled test examples increases by just 6%. On CIFAR-10 with corrected labels: VGG-11 outperforms VGG-19 if the prevalence of originally mislabeled test examples increases by just 5%. Test set errors across the 10 datasets can be viewed at https://labelerrors.com and all label errors can be reproduced by https://github.com/cleanlab/label-errors.

Curtis G. Northcutt, Anish Athalye, Jonas Mueller
arXiv:2103.14749 · stat.ML, cs.AI, cs.LG · submitted Mar 26, 2021 · updated Nov 7, 2021
abstract · pdf · html · Demo available at https://labelerrors.com/ and source code available at https://github.com/cleanlab/label-errors

add comment on HN
Also discussed: Jul 2023 (2 points, 0 comments)

Another somewhat related "issue" is that the mean and std used in ImageNet has unknown origins. It's based off the data but it's not clear exactly how they ended up with those numbers. [0]

[0] https://github.com/pytorch/vision/issues/1439

Interestingly, these are also used for a lot of CV models that have nothing to do with ImageNet.
I wonder if that's because the numbers are general enough across many different tasks or if normalization just isn't as important as we assume?
When I first started playing around with neural networks, I noticed that even the MNIST numbers had some images mislabeled.

I kinda suspect that most people never really check imported datasets for errors.