about
The curious case of the test set AUROC (arxiv.org)
1 point by PaulHoule on Jan 5, 2024 | hide | past | pdf | discuss on HN

In plain words: Models are usually judged by one score from their test data: the area under the ROC curve (hits versus false alarms), or sensitivity and specificity at a validation-picked threshold. Yet such test-curve scores show only a narrow slice of how well a model generalizes.

Abstract

Whilst the size and complexity of ML models have rapidly and significantly increased over the past decade, the methods for assessing their performance have not kept pace. In particular, among the many potential performance metrics, the ML community stubbornly continues to use (a) the area under the receiver operating characteristic curve (AUROC) for a validation and test cohort (distinct from training data) or (b) the sensitivity and specificity for the test data at an optimal threshold determined from the validation ROC. However, we argue that considering scores derived from the test ROC curve alone gives only a narrow insight into how a model performs and its ability to generalise.

Michael Roberts, Alon Hazan, Sören Dittmer, James H. F. Rudd, Carola-Bibiane Schönlieb
arXiv:2312.16188 · cs.LG, stat.ME · submitted Dec 19, 2023
abstract · pdf · html · 3 pages, 4 figures

add comment on HN