In plain words: A newer memory network that mixes past memories works well for text but forgets too quickly for long time series. Feeding it chunks of the series and handling each data stream separately fixed that, beating the best current forecasting models.
Abstract
Traditional recurrent neural network architectures, such as long short-term memory neural networks (LSTM), have historically held a prominent role in time series forecasting (TSF) tasks. While the recently introduced sLSTM for Natural Language Processing (NLP) introduces exponential gating and memory mixing that are beneficial for long term sequential learning, its potential short memory issue is a barrier to applying sLSTM directly in TSF. To address this, we propose a simple yet efficient algorithm named P-sLSTM, which is built upon sLSTM by incorporating patching and channel independence. These modifications substantially enhance sLSTM's performance in TSF, achieving state-of-the-art results. Furthermore, we provide theoretical justifications for our design, and conduct extensive comparative and analytical experiments to fully validate the efficiency and superior performance of our model.
Yaxuan Kong, Zepu Wang, Yuqi Nie, Tian Zhou, Stefan Zohren, Yuxuan Liang, Peng Sun, Qingsong Wen
arXiv:2408.10006 · cs.LG · submitted Aug 19, 2024 · updated Feb 24, 2025
abstract · pdf · html · Accepted by 39th Annual AAAI Conference on Artificial Intelligence (AAAI 2025)