In plain words: Pulling nearby time points together and pushing distant ones apart yields a simple linked chain of representations, so "what next?" and "how did we get here?" reduce to inverting one small matrix. Simulations up to 46 dimensions confirm it, instead of costly sampling over raw data.
Abstract · Inference via Interpolation: Contrastive Representations Provably Enable Planning and Inference
Given time series data, how can we answer questions like "what will happen in the future?" and "how did we get here?" These sorts of probabilistic inference questions are challenging when observations are high-dimensional. In this paper, we show how these questions can have compact, closed form solutions in terms of learned representations. The key idea is to apply a variant of contrastive learning to time series data. Prior work already shows that the representations learned by contrastive learning encode a probability ratio. By extending prior work to show that the marginal distribution over representations is Gaussian, we can then prove that joint distribution of representations is also Gaussian. Taken together, these results show that representations learned via temporal contrastive learning follow a Gauss-Markov chain, a graphical model where inference (e.g., prediction, planning) over representations corresponds to inverting a low-dimensional matrix. In one special case, inferring intermediate representations will be equivalent to interpolating between the learned representations. We validate our theory using numerical simulations on tasks up to 46-dimensions.
Benjamin Eysenbach, Vivek Myers, Ruslan Salakhutdinov, Sergey Levine
arXiv:2403.04082 · cs.LG, stat.ML · submitted Mar 6, 2024 · updated May 21, 2025
abstract · pdf · html · Code: https://github.com/vivekmyers/contrastive_planning