In plain words: Chopping a time series into overlapping chunks before grouping it — the usual trick for making data easier to cluster — can go spectacularly wrong. Testing chunk lengths against series lengths reveals three distinct failure modes, each explained with theory.
Abstract · On the clustering behavior of sliding windows
Things can go spectacularly wrong when clustering timeseries data that has been preprocessed with a sliding window. We highlight three surprising failures that emerge depending on how the window size compares with the timeseries length. In addition to computational examples, we present theoretical explanations for each of these failure modes.
Boris Alexeev, Wenyan Luo, Dustin G. Mixon, Yan X Zhang
arXiv:2503.14393 · cs.LG · submitted Mar 18, 2025
abstract · pdf · html