about
Online Isolation Forest (arxiv.org)
2 points by badmonster on May 15, 2025 | hide | past | pdf | discuss on HN

In plain words: A new anomaly detector built for data streams updates itself as the data changes, instead of storing all past data and retraining from scratch now and then. It matched other streaming detectors, nearly matched offline ones that retrain regularly, and ran fastest of all.

Abstract

The anomaly detection literature is abundant with offline methods, which require repeated access to data in memory, and impose impractical assumptions when applied to a streaming context. Existing online anomaly detection methods also generally fail to address these constraints, resorting to periodic retraining to adapt to the online context. We propose Online-iForest, a novel method explicitly designed for streaming conditions that seamlessly tracks the data generating process as it evolves over time. Experimental validation on real-world datasets demonstrated that Online-iForest is on par with online alternatives and closely rivals state-of-the-art offline anomaly detection techniques that undergo periodic retraining. Notably, Online-iForest consistently outperforms all competitors in terms of efficiency, making it a promising solution in applications where fast identification of anomalies is of primary importance such as cybersecurity, fraud and fault detection.

Filippo Leveni, Guilherme Weigert Cassales, Bernhard Pfahringer, Albert Bifet, Giacomo Boracchi
arXiv:2505.09593 · cs.LG, cs.AI, stat.ML · submitted May 14, 2025
abstract · pdf · html · Accepted at International Conference on Machine Learning (ICML 2024)

add comment on HN