In plain words: It turns an agent's path into a sequence of discrete locations and reads it with a Transformer, the same kind of network behind language models, learning from labels or by filling in hidden steps. The encodings separated different labels and showed which locations are similar.
Abstract
Spatiotemporal data faces many analogous challenges to natural language text including the ordering of locations (words) in a sequence, long range dependencies between locations, and locations having multiple meanings. In this work, we propose a novel model for representing high dimensional spatiotemporal trajectories as sequences of discrete locations and encoding them with a Transformer-based neural network architecture. Similar to language models, our Sequence Transformer for Agent Representation Encodings (STARE) model can learn representations and structure in trajectory data through both supervisory tasks (e.g., classification), and self-supervisory tasks (e.g., masked modelling). We present experimental results on various synthetic and real trajectory datasets and show that our proposed model can learn meaningful encodings that are useful for many downstream tasks including discriminating between labels and indicating similarity between locations. Using these encodings, we also learn relationships between agents and locations present in spatiotemporal data.
Athanasios Tsiligkaridis, Nicholas Kalinowski, Zhongheng Li, Elizabeth Hou
arXiv:2410.09204 · cs.LG, cs.AI, cs.CL · submitted Oct 11, 2024
abstract · pdf · html · 12 pages, to be presented at GeoAI workshop at ACM SigSpatial 2024