In plain words: Instead of treating each timestamp as a token, this model makes each full series one token, so attention learns how different series relate while the network learns each series' own patterns. It beat the usual setup on real-world forecasting, especially with longer history windows.
Abstract · iTransformer: Inverted Transformers Are Effective for Time Series Forecasting
The recent boom of linear forecasting models questions the ongoing passion for architectural modifications of Transformer-based forecasters. These forecasters leverage Transformers to model the global dependencies over temporal tokens of time series, with each token formed by multiple variates of the same timestamp. However, Transformers are challenged in forecasting series with larger lookback windows due to performance degradation and computation explosion. Besides, the embedding for each temporal token fuses multiple variates that represent potential delayed events and distinct physical measurements, which may fail in learning variate-centric representations and result in meaningless attention maps. In this work, we reflect on the competent duties of Transformer components and repurpose the Transformer architecture without any modification to the basic components. We propose iTransformer that simply applies the attention and feed-forward network on the inverted dimensions. Specifically, the time points of individual series are embedded into variate tokens which are utilized by the attention mechanism to capture multivariate correlations; meanwhile, the feed-forward network is applied for each variate token to learn nonlinear representations. The iTransformer model achieves state-of-the-art on challenging real-world datasets, which further empowers the Transformer family with promoted performance, generalization ability across different variates, and better utilization of arbitrary lookback windows, making it a nice alternative as the fundamental backbone of time series forecasting. Code is available at this repository: https://github.com/thuml/iTransformer.
Yong Liu, Tengge Hu, Haoran Zhang, Haixu Wu, Shiyu Wang, Lintao Ma, Mingsheng Long
arXiv:2310.06625 · cs.LG · submitted Oct 10, 2023 · updated Mar 14, 2024
abstract · pdf · html
Let's say you have 100 intersections, and you want to predict the traffic on each in cars/sec. You sample every hour, and you keep 24 hours of context, and try to predict the next 4.
First, you'd make 100 "tokens" (really stretching the meaning of token here), one for each stoplight, and loading 24 samples (the history of that stoplight) into each token, and normalize.
Next, you run each token through a Multi-Layer Perceptron (vanilla, old-school neural network) to make a vector of dim D.
Next, for each layer of the transformer, you: 1. Perform "cross-attention," i.e. the query/key/value dance. This is how the different time series (erm, tokens) get to share information. 2. Normalize across all. 3. Run another bog-standard MLP independently on each token. This is the opportunity to examine the history of each time series. 4. Normalize again across all.
Then, you map each "token" (ugh) from being D-dimensional to 4-dimensional, so for each stoplight it predicts the traffic ahead for the next 4 hours. This is also a regular MLP.
So specifically, if you're only predicting a single time series (one stoplight), this method is equivalent to running a regular neural network.
It also, interestingly enough, skips the cool sinusoidal position embedding that transformers use to embed token position. Fair enough, since here the time dimension is fixed and the index of the feed-forward neurons in each MLP layer corresponds (roughly) to the time index of the sample.
The architecture looks weird to me, but apparently it works so that's cool! But I'm not sure how well it works, and my unscientific gut feel is that there's a better and simpler architecture crying out to be found, because this looks a bit tortured. Like, nothing in it explicitly models the time dimension - that task is left to the MLPs - and that seems weird.