In plain words: They checked the popular anomaly-detection benchmark datasets example by example and found most have one of four flaws. That makes past algorithm rankings unreliable, so they released a cleaner archive for fair comparisons.
Abstract · Current Time Series Anomaly Detection Benchmarks are Flawed and are Creating the Illusion of Progress
Time series anomaly detection has been a perennially important topic in data science, with papers dating back to the 1950s. However, in recent years there has been an explosion of interest in this topic, much of it driven by the success of deep learning in other domains and for other time series tasks. Most of these papers test on one or more of a handful of popular benchmark datasets, created by Yahoo, Numenta, NASA, etc. In this work we make a surprising claim. The majority of the individual exemplars in these datasets suffer from one or more of four flaws. Because of these four flaws, we believe that many published comparisons of anomaly detection algorithms may be unreliable, and more importantly, much of the apparent progress in recent years may be illusionary. In addition to demonstrating these claims, with this paper we introduce the UCR Time Series Anomaly Archive. We believe that this resource will perform a similar role as the UCR Time Series Classification Archive, by providing the community with a benchmark that allows meaningful comparisons between approaches and a meaningful gauge of overall progress.
Renjie Wu, Eamonn J. Keogh
arXiv:2009.13807 · cs.LG, stat.ML · submitted Sep 29, 2020 · updated Sep 3, 2022
abstract · pdf · Full paper accepted by IEEE TKDE, extended abstract accepted by IEEE ICDE 2022
The following quote from the article is striking: "Qiu et al. introduce a “novel anomaly detector for time-series KPIs based on supervised deep-learning models with convolution and long short-term memory (LSTM) neural networks, and a variational auto-encoder (VAE) oversampling model.” This description sounds like it has many “moving parts”, and indeed, the dozen or so explicitly listed parameters include: convolution filter, activation, kernel size, strides, padding, LSTM input size, dense input size, softmax loss function, window size, learning rate and batch size. All of this is to demonstrate “accuracy exceeding 0.90.” However, as we will show, much of the results of this complex approach can be duplicated with a single line of code and a few minutes of effort."