about
The influence of random seeds in deep learning architectures for computer vision (arxiv.org)
3 points by PaulHoule on May 16, 2023 | hide | past | pdf | 1 comment on HN

In plain words: Tested up to 10,000 random starting values for training image-recognition models on two datasets to see how much the choice changes accuracy. Scores barely shift on average, yet a seed that does much better or worse than usual is surprisingly easy to find.

Abstract · Torch.manual_seed(3407) is all you need: On the influence of random seeds in deep learning architectures for computer vision

In this paper I investigate the effect of random seed selection on the accuracy when using popular deep learning architectures for computer vision. I scan a large amount of seeds (up to $10^4$) on CIFAR 10 and I also scan fewer seeds on Imagenet using pre-trained models to investigate large scale datasets. The conclusions are that even if the variance is not very large, it is surprisingly easy to find an outlier that performs much better or much worse than the average.

David Picard
arXiv:2109.08203 · cs.CV · submitted Sep 16, 2021 · updated May 11, 2023
abstract · pdf · html · fixed typos

add comment on HN
Also discussed: Oct 2024 (2 points, 0 comments) · May 2023 (1 point, 0 comments)

It certainly vexes me that, when training and evaluating models, I sometimes get better results and sometimes worse. There's a certain amount of randomness associated with the evaluation data (e.g. you might get the right or wrong answer for a particular question with a binomial distribution, or in a real world problem where you train + eval everyday you'll find some days you have easier problems than other days.) Another problem though is that you might sometimes get a better model or sometimes get a worse model and that's what this paper is about.

It drives me up the wall that people write papers where they compare a large number of models and mark the highest accuracy/F1/AUC in bold but never do an analysis to ask "are the differences significant?"