about
Anomaly Detection: How to Artificially Increase Your F1-Score (arxiv.org)
1 point by belter on Jul 1, 2021 | hide | past | pdf | 1 comment on HN

In plain words: They tested how anomaly-detection scores change when the share of odd cases or the train-test split is altered. F1 and precision-recall scores rose just from tweaking that split, so they recommend a ranking score instead, which stays steadier.

Abstract · Anomaly Detection: How to Artificially Increase your F1-Score with a Biased Evaluation Protocol

Anomaly detection is a widely explored domain in machine learning. Many models are proposed in the literature, and compared through different metrics measured on various datasets. The most popular metrics used to compare performances are F1-score, AUC and AVPR. In this paper, we show that F1-score and AVPR are highly sensitive to the contamination rate. One consequence is that it is possible to artificially increase their values by modifying the train-test split procedure. This leads to misleading comparisons between algorithms in the literature, especially when the evaluation protocol is not well detailed. Moreover, we show that the F1-score and the AVPR cannot be used to compare performances on different datasets as they do not reflect the intrinsic difficulty of modeling such data. Based on these observations, we claim that F1-score and AVPR should not be used as metrics for anomaly detection. We recommend a generic evaluation procedure for unsupervised anomaly detection, including the use of other metrics such as the AUC, which are more robust to arbitrary choices in the evaluation protocol.

Damien Fourure, Muhammad Usama Javaid, Nicolas Posocco, Simon Tihon
arXiv:2106.16020 · cs.LG · submitted Jun 30, 2021
abstract · pdf · html · 16 pages, 7 figures, to be published in ECML-PKDD 2021, for official implementation see https://github.com/euranova/F1-Score-is-Biased

add comment on HN

"Anomaly detection is a widely explored domain in machine learning. Many models are proposed in the literature, and compared through different metrics measured on various datasets. The most popular metrics used to compare performances are F1-score, AUC and AVPR. In this paper, we show that F1-score and AVPR are highly sensitive to the contamination rate. One consequence is that it is possible to artificially increase their values by modifying the train-test split procedure. This leads to misleading comparisons between algorithms in the literature,especially when the evaluation protocol is not well detailed. Moreover, we show that the F1-score and the AVPR cannot be used to compare performances on different datasets as they do not reflect the intrinsic difficulty of modeling such data. Based on these observations, we claim that F1-score and AVPR should not be used as metrics for anomaly detection. We recommend a generic evaluation procedure for unsupervised anomaly detection, including the use of other metrics such as the AUC, which are more robust to arbitrary choices in the evaluation protocol."

https://arxiv.org/pdf/2106.16020.pdf