In plain words: Large forecasting models trained on many kinds of data were tested on cloud data without extra training on it. Simple straight-line baselines beat them every time, and some model forecasts suddenly looked random and erratic.
Abstract
Time series foundation models (FMs) have emerged as a popular paradigm for zero-shot multi-domain forecasting. FMs are trained on numerous diverse datasets and claim to be effective forecasters across multiple different time series domains, including cloud data. In this work we investigate this claim, exploring the effectiveness of FMs on cloud data. We demonstrate that many well-known FMs fail to generate meaningful or accurate zero-shot forecasts in this setting. We support this claim empirically, showing that FMs are outperformed consistently by simple linear baselines. We also illustrate a number of interesting pathologies, including instances where FMs suddenly output seemingly erratic, random-looking forecasts. Our results suggest a widespread failure of FMs to model cloud data.
William Toner, Thomas L. Lee, Artjom Joosen, Rajkarn Singh, Martin Asenov
arXiv:2502.12944 · cs.LG · submitted Feb 18, 2025 · updated May 19, 2025
abstract · pdf · html · 5 pages, presented at the "I Can't Believe It's Not Better" workshop at ICLR 2025